Storage and versioning
A document in ADM is metadata plus one or more versions, and a version is what actually has bytes. Where those bytes live — a database BLOB or an OCI Object Storage bucket — is a setting, and nothing above the storage layer needs to know which it is.
Documents and versions
Section titled “Documents and versions”adm_documents one row per document: name, folder, owner, retention, flags └── adm_document_versions one row per version: size, mime type, checksum, contentUploading a file whose name already exists in that folder does not overwrite anything — it adds a version, and the newest one becomes current. So:
- Every upload is recoverable. Somebody replacing a good file with a bad one is undone by restoring the previous version, not by going to a backup.
- Storage grows with every upload. A document with twenty versions holds twenty copies of the content. Versions can be deleted individually.
- The file’s identity — its id, its shares, its tags, its comments, its audit history — belongs to the document, and survives across versions. A share does not have to be re-issued because somebody uploaded a new revision.
Each version carries a checksum of its content, computed on upload. Comparing the stored content against it detects a file that changed underneath ADM, which is a risk when content lives outside the database.
→ Working with documents for the day-to-day operations: version history, restoring, comparing, deleting a version.
Where content lives
Section titled “Where content lives”Two locations, one API
Section titled “Two locations, one API”| Location | Content column | Suits |
|---|---|---|
DATABASE | A BLOB in adm_document_versions | Simplicity. One backup covers everything, no external dependency, no network in the download path. |
OBJECT_STORAGE | An object in an OCI bucket, referenced by key | Volume and cost. Keeps the database small, and lets cold content sit in a cheaper tier. |
Your own code should never care which it is. Read content through
adm_storage_api.get_file_content,
which resolves the location and returns a BLOB either way:
declare l_content blob;begin l_content := adm_storage_api.get_file_content(p_version_id => l_version_id);end;/Choosing per folder, not just globally
Section titled “Choosing per folder, not just globally”Object storage is switched on globally with the OBJECT_STORAGE_ENABLED setting, but the decision
can be refined per folder path with storage policies:
-- everything under /groups/archive/ goes to object storage,-- regardless of the global defaultdeclare l_policy_id number;begin adm_context_api.system_login;
l_policy_id := adm_storage_api.create_storage_policy( p_folder_path_pattern => '/groups/archive/' , p_storage_location => 'OBJECT_STORAGE' , p_description => 'Cold archive' );
commit;end;/Policies can be activated, deactivated and deleted, and
adm_storage_api.determine_storage_location(p_folder_path) tells you which location a given path
resolves to. Check a policy with it before you upload a terabyte through it.
Moving content between locations
Section titled “Moving content between locations”Nothing has to be re-uploaded to change location. The migration procedures move existing content in the background, per document or per folder subtree, in batches:
| Procedure | Moves |
|---|---|
migrate_to_object_storage / migrate_to_blob_storage | One document version. |
migrate_folder_to_object_storage / migrate_folder_to_blob_storage | A folder subtree. |
apply_adm_object_storage_policies | Everything, to wherever the policies say it belongs. |
Migrations are not instantaneous: work is queued and the daily job picks it up, retrying failures up
to OBJECT_STORAGE_RETRY_ATTEMPTS times and processing MIGRATION_BATCH_SIZE files per run.
→ Object storage setup for credentials, buckets and monitoring a migration · Jobs and maintenance
Delayed deletion
Section titled “Delayed deletion”When a version whose content is in object storage is deleted, the object is not removed
immediately — the deletion is scheduled, with a reason
(DOCUMENT_DELETED, VERSION_DELETED, MIGRATION_TO_DATABASE, MANUAL_CLEANUP) and an optional
delay in days, and the daily job carries it out later. A scheduled deletion can be cancelled before
it runs.
The delay covers the case where the database says a file is gone but you need it back: the row is deleted, the bytes are not, for as long as the delay lasts.
Checksum verification
Section titled “Checksum verification”Content in a bucket can change or disappear without the database knowing: someone with bucket access
deletes an object, a lifecycle rule archives it, a migration half-fails. So ADM re-reads objects and
compares them against the checksum recorded at upload time, spread over time rather than all at once:
CHECKSUM_VERIFICATION_BATCH_SIZE files per daily run, and each file re-verified every
CHECKSUM_REVERIFICATION_DAYS days. Mismatches are logged, not silently corrected.
Turn it off with OBJECT_STORAGE_VERIFY_CHECKSUMS if you have equivalent guarantees elsewhere.
Without it, a corrupted file is discovered by the user who needs it.
Extracted metadata
Section titled “Extracted metadata”On upload — and on every new version — ADM reads technical metadata out of the file itself and
stores it as annotations prefixed file.: title, author, created and modified dates, page count,
word count, slide and sheet counts, sheet names, and image dimensions. Supported formats are OOXML
office documents, PDF (best effort — an encrypted PDF may yield nothing) and PNG/JPEG/GIF images.
Extraction is best-effort: it never breaks an upload, and it never raises. If a file yields nothing,
the upload still succeeds. file.extracted_at is always written, which is how the backfill job knows
not to look at that document again.
The mime type stored is the one derived from the file content, not the one the browser reported, because the reported type is not trustworthy.
Metadata extraction is controlled by the FILE_METADATA_EXTRACTION setting. Existing documents from
before it was enabled are picked up by the backfill, which processes
FILE_METADATA_BACKFILL_BATCH_SIZE documents per daily run.
The extracted values are ordinary annotations, so they are queryable:
select doc.document_name , ann.annotation_value as page_count from adm_documents doc join adm_document_annotations ann on ann.document_id = doc.document_id where ann.annotation_key = 'file.page_count' and to_number(ann.annotation_value) > 100;