Skip to content

Storage and versioning

A document in ADM is metadata plus one or more versions, and a version is what actually has bytes. Where those bytes live — a database BLOB or an OCI Object Storage bucket — is a setting, and nothing above the storage layer needs to know which it is.

adm_documents one row per document: name, folder, owner, retention, flags
└── adm_document_versions one row per version: size, mime type, checksum, content

Uploading a file whose name already exists in that folder does not overwrite anything — it adds a version, and the newest one becomes current. So:

  • Every upload is recoverable. Somebody replacing a good file with a bad one is undone by restoring the previous version, not by going to a backup.
  • Storage grows with every upload. A document with twenty versions holds twenty copies of the content. Versions can be deleted individually.
  • The file’s identity — its id, its shares, its tags, its comments, its audit history — belongs to the document, and survives across versions. A share does not have to be re-issued because somebody uploaded a new revision.

Each version carries a checksum of its content, computed on upload. Comparing the stored content against it detects a file that changed underneath ADM, which is a risk when content lives outside the database.

→ Working with documents for the day-to-day operations: version history, restoring, comparing, deleting a version.

LocationContent columnSuits
DATABASEA BLOB in adm_document_versionsSimplicity. One backup covers everything, no external dependency, no network in the download path.
OBJECT_STORAGEAn object in an OCI bucket, referenced by keyVolume and cost. Keeps the database small, and lets cold content sit in a cheaper tier.

Your own code should never care which it is. Read content through adm_storage_api.get_file_content, which resolves the location and returns a BLOB either way:

declare
l_content blob;
begin
l_content := adm_storage_api.get_file_content(p_version_id => l_version_id);
end;
/

Object storage is switched on globally with the OBJECT_STORAGE_ENABLED setting, but the decision can be refined per folder path with storage policies:

-- everything under /groups/archive/ goes to object storage,
-- regardless of the global default
declare
l_policy_id number;
begin
adm_context_api.system_login;
l_policy_id := adm_storage_api.create_storage_policy(
p_folder_path_pattern => '/groups/archive/'
, p_storage_location => 'OBJECT_STORAGE'
, p_description => 'Cold archive'
);
commit;
end;
/

Policies can be activated, deactivated and deleted, and adm_storage_api.determine_storage_location(p_folder_path) tells you which location a given path resolves to. Check a policy with it before you upload a terabyte through it.

Nothing has to be re-uploaded to change location. The migration procedures move existing content in the background, per document or per folder subtree, in batches:

ProcedureMoves
migrate_to_object_storage / migrate_to_blob_storageOne document version.
migrate_folder_to_object_storage / migrate_folder_to_blob_storageA folder subtree.
apply_adm_object_storage_policiesEverything, to wherever the policies say it belongs.

Migrations are not instantaneous: work is queued and the daily job picks it up, retrying failures up to OBJECT_STORAGE_RETRY_ATTEMPTS times and processing MIGRATION_BATCH_SIZE files per run.

→ Object storage setup for credentials, buckets and monitoring a migration · Jobs and maintenance

When a version whose content is in object storage is deleted, the object is not removed immediately — the deletion is scheduled, with a reason (DOCUMENT_DELETED, VERSION_DELETED, MIGRATION_TO_DATABASE, MANUAL_CLEANUP) and an optional delay in days, and the daily job carries it out later. A scheduled deletion can be cancelled before it runs.

The delay covers the case where the database says a file is gone but you need it back: the row is deleted, the bytes are not, for as long as the delay lasts.

Content in a bucket can change or disappear without the database knowing: someone with bucket access deletes an object, a lifecycle rule archives it, a migration half-fails. So ADM re-reads objects and compares them against the checksum recorded at upload time, spread over time rather than all at once: CHECKSUM_VERIFICATION_BATCH_SIZE files per daily run, and each file re-verified every CHECKSUM_REVERIFICATION_DAYS days. Mismatches are logged, not silently corrected.

Turn it off with OBJECT_STORAGE_VERIFY_CHECKSUMS if you have equivalent guarantees elsewhere. Without it, a corrupted file is discovered by the user who needs it.

→ Settings reference

On upload — and on every new version — ADM reads technical metadata out of the file itself and stores it as annotations prefixed file.: title, author, created and modified dates, page count, word count, slide and sheet counts, sheet names, and image dimensions. Supported formats are OOXML office documents, PDF (best effort — an encrypted PDF may yield nothing) and PNG/JPEG/GIF images.

Extraction is best-effort: it never breaks an upload, and it never raises. If a file yields nothing, the upload still succeeds. file.extracted_at is always written, which is how the backfill job knows not to look at that document again.

The mime type stored is the one derived from the file content, not the one the browser reported, because the reported type is not trustworthy.

Metadata extraction is controlled by the FILE_METADATA_EXTRACTION setting. Existing documents from before it was enabled are picked up by the backfill, which processes FILE_METADATA_BACKFILL_BATCH_SIZE documents per daily run.

The extracted values are ordinary annotations, so they are queryable:

select doc.document_name
, ann.annotation_value as page_count
from adm_documents doc
join adm_document_annotations ann on ann.document_id = doc.document_id
where ann.annotation_key = 'file.page_count'
and to_number(ann.annotation_value) > 100;