Skip to content

adm_file_metadata_api

Extracts technical file metadata (title, author, page count, image dimensions, …) from uploaded files and stores it as document annotations with the reserved key prefix “file.”.

Supported formats: OOXML office documents (docx/xlsx/pptx incl. macro variants), PDF (best effort - encrypted or compressed-xref PDFs may yield nothing) and images (PNG, JPEG, GIF).

Extraction is best effort: it must never break an upload. The annotation “file.extracted_at” is always written as a sentinel so documents are not re-processed by the backfill.

Backfill metadata for existing documents that have never been processed (no “file.extracted_at” annotation). Fetches the latest version content through adm_storage_api.get_file_content, so object storage files are downloaded - use p_max_documents to batch large instances.

Requires an established context (e.g. adm_context_api.system_login). Caller controls the transaction (no commit inside). Returns 0 immediately when the FILE_METADATA_EXTRACTION setting is disabled.

Signature:

function backfill_metadata (
p_max_documents in number default null
) return number;

Parameters:

NameDirectionTypeDescription
p_max_documentsinnumber default nullOptional cap of documents to process per call

Returns: number - Number of documents processed