Skip to content

Creating RAG configurations

A RAG configuration indexes and embeds a subset of your documents so they can be searched by vector similarity and used to answer questions.

The AI Pack has to be installed first — see Installing the AI Pack for the uc_ai dependency, install_ai.sql, and the web credentials. The five AI settings are listed in the settings reference.

Embeddings can be stored in one of two vector backends, chosen per collection with the vector_store field:

  • ORACLE — native Oracle AI Vector Search, available on Oracle 23ai / 26ai. Nothing external to run; no Qdrant settings needed.
  • QDRANT — the external, free and open-source Qdrant vector database (hostable via its Docker image). Use this on databases without native vector support.

AI_CHUNK_FILE_LIMIT (limit of files chunked/embedded per run) applies to both. When using Qdrant, also configure:

  • AI_QDRANT_COLLECTION_PREFIX: Prefix for Qdrant collection names
  • AI_QDRANT_URL: URL for Qdrant service
  • AI_QDRANT_WEB_CREDENTIAL_ID: APEX Web Credential Static ID for Qdrant authorization

Additionally you need an LLM provider for embedding — and optionally for answer generation, query rewriting, and LLM text extraction — configured via APEX Web Credentials.

A RAG collection enables retrieval augmented generation for a subset of files in ADM. Each collection has its own configuration.

declare
l_collection_id number;
l_config_json clob;
begin
l_config_json := '{"source_folder_path": ... see next chapter}';
l_collection_id := adm_ai_rag_api.create_rag_collection(
p_collection_name => 'test_collection',
p_description => 'Test RAG collection',
p_config_json => l_config_json
);
end;

To choose which files are included, pass either a folder path or an annotation key (any file carrying that annotation, or any file in a folder carrying it). All of those files are then chunked and embedded automatically.

A collection’s configuration is a single JSON object. A few sections are required; the rest are optional add-ons you can layer on as needed:

Section / fieldRequiredPurpose
source_folder_path or source_file_annotation_keyYesWhich files belong to the collection
content_typeYesKind of content (currently "text")
embeddingYesHow chunks are turned into vectors
chunksYesHow documents are split into chunks
distance_metric, vector_store, default_*NoRetrieval tuning (root-level fields)
answer_modelNoLLM used by generate_answer
query_rewritingNoRewrite the user query before searching
llm_text_extractionNoExtract file text with an LLM (multimodal file input) instead of the built-in filter — for messy or scanned PDFs
hybrid_searchNoBlend Oracle Text keyword search with vector search (RRF) so exact-term matches are not missed

The example below shows every section together; the reference tables that follow explain each field.

{
"source_folder_path": "/users/philipp/rag",
"source_file_annotation_key": "jira-articles",
"content_type": "text",
"distance_metric": "Cosine",
"default_score_threshold": 0.5,
"default_limit": 20,
"default_max_sources": 5,
"default_surrounding_chunks": 1,
"embedding": {
"dimensions": 1536,
"provider": "openai",
"model": "text-embedding-3-small",
"config": {
"g_apex_web_credential": "OPENAI_API_KEY"
}
},
"chunks": {
"target_chunk_size": 2048,
"overlap_size": 256,
"min_chunk_size": 512,
"max_chunk_size": 2560
},
"answer_model": {
"provider": "openai",
"model": "gpt-4o",
"system_prompt": "You are a support assistant. Answer using only the provided sources and cite them.",
"config": {
"g_apex_web_credential": "OPENAI_API_KEY",
"g_enable_reasoning": true,
"g_reasoning_level": "low",
"openai": {
"g_reasoning_effort": "low"
}
}
},
"query_rewriting": {
"enabled": true,
"provider": "openai",
"model": "gpt-4o",
"config": {
"g_apex_web_credential": "OPENAI_API_KEY"
}
},
"llm_text_extraction": {
"provider": "google",
"model": "gemini-2.5-flash",
"apply_to": {
"mime_types": ["application/pdf"]
},
"config": {
"g_apex_web_credential": "GEMINI_API_KEY"
}
},
"hybrid_search": {
"enabled": true,
"rrf_k": 60
}
}

Each AI section (embedding, answer_model, query_rewriting) keeps its provider and model as explicit fields, and passes everything else through a generic config object straight to the UC AI SDK. The keys map directly onto UC AI’s settings — common keys at the root and provider-specific keys under a nested provider object:

KeyLevelDescription
g_apex_web_credentialRootAPEX web credential (API key) for the provider. Replaces the old provider_web_credential.
g_base_urlRootBase URL override for the API endpoint. Replaces the old provider_base_url_override.
g_enable_reasoningRoottrue/false — enable reasoning/thinking for models that support it.
g_reasoning_levelRoot"low", "medium" or "high". Replaces the old reasoning field.
openai.g_reasoning_effortProviderOpenAI reasoning effort ("minimal", "low", "medium", "high", "xhigh").
anthropic.g_max_tokensProviderAnthropic max output tokens.
google.g_embedding_output_dimensionsProviderGoogle embedding output dimensions.

Anything UC AI supports can be set here — see the UC AI documentation for the full list of root and per-provider (openai, anthropic, google, ollama, oci, xai, openrouter) settings. The keys are validated by UC AI when a call is made, so an unknown key raises an error at generation time. config is optional and may be omitted entirely.

FieldRequiredTypeDescription
source_folder_pathConditionalStringThe folder path within your document management system to index. Either this or source_file_annotation_key must be provided.
source_file_annotation_keyConditionalStringAn annotation key (either file or folder annotations) to filter which files should be included in the RAG collection. Either this or source_folder_path must be provided.
content_typeYesStringThe type of content being indexed. Currently only "text" is supported.
distance_metricNoStringThe distance metric used for vector similarity search. Defaults to "Cosine". Other options may include "Euclidean" or "Dot".
vector_storeNoStringWhere embeddings are stored/searched: "ORACLE" (native Oracle AI Vector Search, requires Oracle 23ai/26ai) or "QDRANT" (external Qdrant server). Materialized on create ("ORACLE" on 23ai+, otherwise "QDRANT") and cannot be changed afterwards.
default_score_thresholdNoNumberFallback minimum similarity score (0–1) used by search/answer calls when no p_score_threshold argument is passed. If omitted it is materialized to 0.5 on create so you can tune it later. See the note below on model-dependent scales.
default_limitNoNumberFallback vector search top-k (how many chunks the similarity search returns) used when no p_limit argument is passed. If omitted it is materialized to 20 on create. A larger pool improves source diversity; the score threshold still filters weak matches.
default_max_sourcesNoNumberFallback maximum number of source documents returned by prepare_sources/generate_answer when no p_max_sources argument is passed. Documents are ranked by their best chunk score, so this returns the most relevant ones. If omitted it is materialized to 5 on create.
default_surrounding_chunksNoNumberFallback number of neighboring chunks included around each matched chunk (see the note below). 0 disables expansion. If omitted it is materialized to 1 on create.

These settings configure how document chunks are converted to vector embeddings.

FieldRequiredTypeDescription
dimensionsYesNumberThe number of dimensions for the embedding vectors. Must match the output dimensions of your chosen embedding model (e.g., 1536 for OpenAI’s text-embedding-3-small).
providerYesStringThe AI provider for generating embeddings (e.g., "openai", "cohere", "ollama").
modelYesStringThe specific embedding model to use (e.g., "text-embedding-3-small", "text-embedding-3-large").
configNoObjectGeneric UC AI config passthrough (see The generic config object). Set the API key with g_apex_web_credential and, if needed, g_base_url.

These settings control how documents are split into smaller chunks for processing.

FieldRequiredTypeDescription
target_chunk_sizeYesNumberThe target size (in characters) for each chunk. Recommended: 1024-2048.
overlap_sizeYesNumberThe number of characters that overlap between consecutive chunks. Helps preserve context. Recommended: 128-256.
min_chunk_sizeYesNumberThe minimum allowed chunk size. Chunks smaller than this may be merged with adjacent chunks.
max_chunk_sizeYesNumberThe maximum allowed chunk size. Chunks larger than this will be split further.

These settings configure the AI model used to generate answers from retrieved context. This section is optional but required if you want to use the generate_answer functionality.

FieldRequiredTypeDescription
providerYes (if section present)StringThe AI provider for generating answers (e.g., "openai", "anthropic").
modelYes (if section present)StringThe specific model to use for answer generation (e.g., "gpt-4o", "gpt-4o-mini").
system_promptNoStringCustom system prompt used for answer generation when no p_system_prompt argument is passed. If omitted, a built-in prompt is used (answer only from the sources, cite them).
configNoObjectGeneric UC AI config passthrough (see The generic config object). Set the API key with g_apex_web_credential, reasoning with g_enable_reasoning + g_reasoning_level, etc.

Query Rewriting Configuration (query_rewriting)

Section titled “Query Rewriting Configuration (query_rewriting)”

These settings configure optional query rewriting to improve search results. The entire section is optional.

FieldRequiredTypeDescription
enabledNoBooleanWhether to enable query rewriting. Set to true to activate.
providerYes (if enabled)StringThe AI provider for query rewriting.
modelYes (if enabled)StringThe model to use for rewriting queries.
configNoObjectGeneric UC AI config passthrough (see The generic config object). Set the API key with g_apex_web_credential, etc.

Text Extraction Configuration (llm_text_extraction)

Section titled “Text Extraction Configuration (llm_text_extraction)”

By default ADM extracts a document’s text with Oracle Text’s built-in filter (the same engine used for full-text search) and stores it as the searchable content that chunks are built from. For complex or scanned PDFs this can be noisy — repeated page headers/footers, table-of-contents dot leaders, broken tables — which hurts retrieval quality.

This optional section instead sends matching files to a multimodal LLM (via file input) and stores the model’s clean transcription. It is an opt-in override: the built-in extractor still handles every file the section does not target, so adding this section never changes extraction for other file types.

FieldRequiredTypeDescription
providerYes (if section present)StringUC AI provider whose model accepts file input (e.g. "google", "anthropic", "openai", or "oci" with a file-capable model).
modelYes (if section present)StringModel used to transcribe the file (e.g. "gemini-2.5-flash"). Must accept PDF/file input.
configNoObjectGeneric UC AI config passthrough (see The generic config object). Set the API key with g_apex_web_credential. For large documents raise the provider’s max output tokens (e.g. oci.g_max_tokens) so the transcription is not truncated.
promptNoStringOverride the default extraction prompt (the default asks for a clean, complete transcription and to drop page headers/footers and TOC leaders).
apply_toNoObjectWhich files use the LLM. Defaults to { "mime_types": ["application/pdf"] }.
apply_to.mime_typesNoArrayOnly these MIME types are sent to the LLM; every other type keeps the built-in extractor. Defaults to ["application/pdf"].
apply_to.annotation_keyNoStringRestrict further to documents carrying this ADM annotation — tag just the known-problem files to keep LLM cost down. Combined with mime_types (AND).
apply_to.annotation_valueNoStringOptional value the annotation must equal (requires annotation_key).
Section titled “Hybrid Search Configuration (hybrid_search)”

By default retrieval is pure semantic vector search. Vector search can miss lexical / near-verbatim matches — exact terms, part numbers, standard codes, proper names — when their embedding similarity happens to be low. This optional section runs an Oracle Text keyword search alongside the vector search and fuses the two result lists with Reciprocal Rank Fusion (RRF), so a chunk that strongly matches the query’s words ranks highly even if its vector score is weak (and vice-versa). Absent this section, retrieval is vector-only (unchanged).

It reuses the Oracle Text index ADM already maintains on the extracted document text and maps keyword hits back to the exact chunk, so there is no extra schema and it works whether the collection’s vectors live in Oracle or Qdrant.

FieldRequiredTypeDescription
enabledYes (if section present)BooleanSet to true to turn hybrid retrieval on.
rrf_kNoNumberRRF constant (default 60). Higher = flatter fusion (rank differences matter less); lower = the very top ranks dominate.
keyword_doc_candidatesNoNumberHow many top keyword-matching documents to inspect for chunk-level hits (default 20).
vector_weightNoNumberWeight of the vector result list in the fusion (default 1).
keyword_weightNoNumberWeight of the keyword result list in the fusion (default 1). Raise above 1 to favor exact-term (lexical) matches.

Oracle vector store only — ignored when vector_store is "QDRANT".

Every key is optional. Left out entirely, a collection gets an IVF index at target accuracy 95, unpartitioned and unquantized, which is what the two materialized keys (type and target_accuracy) record on create so you can see and edit them.

The index is not fixed at creation. Edit this section with update_rag_collection, then run the reconcile:

begin
adm_ai_vector_api.apply_index_config(p_rag_collection_id => 42);
end;
/

apply_index_config compares the configuration against what was last applied and does the least work that makes them agree — nothing, a drop-and-recreate of the index, or a rebuild of the store table when the partitioning changed. It is idempotent, so running it twice costs nothing, and it is the same call a release migration makes.

Like vector_store and the retrieval defaults, this section is carried over when an update_rag_collection config omits it. An edit to some unrelated part of the configuration will not reset the index back to the defaults and send the next reconcile off to rebuild the store.

FieldRequiredTypeDescription
typeNoString"IVF" (default), "HNSW", "AUTO" or "NONE". See the note below on choosing.
target_accuracyNoNumberBuild-time target accuracy, 1–100. Defaults to 95. Oracle derives the internal parameters from it, which is why they are all optional.
parallel_degreeNoNumberDegree of parallelism for building the index, 1–1024. Omitted by default.
hnsw.neighborsNoNumberMaximum connections per vector, 2–2048. Omit to let Oracle derive it (its own fallback is 32).
hnsw.efconstructionNoNumberCandidates considered per insertion, 1–65535. Omit to let Oracle derive it (its own fallback is 300).
ivf.neighbor_partitionsNoNumberTarget number of centroid partitions, 1–10000000. Omit to let Oracle size it from the row count.
ivf.samples_per_partitionNoNumberVectors passed to the clustering algorithm per partition.
ivf.min_vectors_per_partitionNoNumberTrims partitions smaller than this. 0 disables trimming.
quantization.algorithmNoString"NONE" (default) or "SCALAR". HNSW only — rejected for IVF.
quantization.compression_ratioNoNumber2, 4 or 8. Defaults to 4 when quantization is on.
quantization.rescore_factorNoNumber1–100. Rescores quantized candidates against the full vectors to recover accuracy.
partitioning.enabledNoBooleanfalse by default. Hash partitions the store on the owning file and builds a LOCAL index. IVF only — Oracle does not support local HNSW indexes.
partitioning.partition_countNoNumberNumber of hash partitions, 1–1024. Defaults to 8 when partitioning is on.
search.target_accuracyNoNumberQuery-time accuracy override, 1–100. Absent means “search at the accuracy the index was built for”.
search.efsearchNoNumberHNSW candidates to consider per query, 1–65535. Raise it above efconstruction for more accurate results.
search.neighbor_partition_probesNoNumberIVF partitions to probe per query. Raise it for more accurate results.
"vector_index": {
"type": "IVF",
"target_accuracy": 95,
"ivf": { "neighbor_partitions": 100 },
"partitioning": { "enabled": true, "partition_count": 8 },
"search": { "neighbor_partition_probes": 5 }
}

Here is the minimal configuration required to create a RAG collection:

{
"source_folder_path": "/documents/my-folder",
"content_type": "text",
"embedding": {
"dimensions": 1536,
"provider": "openai",
"model": "text-embedding-3-small",
"config": {
"g_apex_web_credential": "OPENAI_API_KEY"
}
},
"chunks": {
"target_chunk_size": 2048,
"overlap_size": 256,
"min_chunk_size": 512,
"max_chunk_size": 2560
}
}
  • Embedding Dimensions: Always verify the output dimensions of your chosen embedding model. Using incorrect dimensions will cause errors.
  • Chunk Sizes: Larger chunks provide more context but may reduce precision. Smaller chunks are more precise but may lose context. Start with the recommended values and adjust based on your use case.
  • Overlap: The overlap helps ensure that important information split across chunk boundaries is still captured. A value of 10-15% of the target chunk size is typical.
  • Web Credentials: Create your APEX web credentials before setting up the RAG configuration. The credential names are case-sensitive.

Two scheduled jobs keep a collection current, and they run at different rates because they cost very different things.

Discovery — ADM_AI_HOURLY_JOB, hourly. Works out which documents match each collection’s criteria, adds a file link for the ones that have joined, and retires the ones that have left. This is the expensive half: it re-evaluates documents against every collection’s configuration.

Processing — ADM_AI_RAG_QUEUE_JOB, every five minutes. Extracts the text of the files discovery found, chunks it, embeds the chunks, and drops the vectors of files that have left. So a document that already sits in a collection’s folder is normally searchable within minutes of being uploaded; one that has to be discovered first waits for the next hourly run.

Three settings pace the processing job. The defaults suit most installations:

SettingDefaultWhat it does
AI_JOB_BATCH_SIZE10Files or jobs per batch, and the commit unit
AI_JOB_MAX_SECONDS240Budget for one run; keep it below the job interval
AI_JOB_MAX_ATTEMPTS3Attempts before a job or a file is written off

A run that uses up its budget stops starting batches; what it did not reach stays queued for the next run five minutes later. If you want more done per run rather than more runs, raise AI_JOB_MAX_SECONDS and the job interval together — AI_JOB_MAX_SECONDS above the interval means runs overlap.

To make a collection catch up immediately rather than waiting for the next job — after changing its configuration, for instance — run the whole pipeline for it once:

begin
adm_ai_rag_workflow_api.rag_collection_maintenance (
p_collection_id => 1
);
end;
/

That is what the Sync Collection button on the RAG collection page does. Discovery on its own is sync_rag_collection_files, which only decides membership and queues the work:

begin
adm_ai_rag_workflow_api.sync_rag_collection_files (
p_collection_id => 1
);
end;
/

Neither call conflicts with the scheduled job: each unit of work is claimed by whichever run reaches it first, and the other skips it.

adm_ai_rag_collection_files_v gives a per-file indexing_status — PENDING, PROCESSING, INCOMPLETE, INDEXED, or ERROR with the reason in latest_error_message. A file that is failing but still inside its attempt budget shows as PENDING; it becomes ERROR once the budget is used up.

For the queue itself, adm_ai_rag_job_queue_v carries minutes_since_created, which is the lag to watch:

select job_type
, status
, count(*) as jobs
, max(minutes_since_created) as oldest_minutes
from adm_ai_rag_job_queue_v
where status in ('PENDING', 'IN_PROGRESS')
group by job_type, status;

What each run did is in adm_job_logs under the job id process_rag_queues, and adm_job_status records its last run and last error.

Once your RAG collection is set up and synchronized, you can query it programmatically using the adm_ai_rag_api package.

Use search_chunks to find document chunks that match a query. This returns individual chunks with their similarity scores:

select rag_chunk_id
, chunk_text
, score
, document_name
, version_number
from table(adm_ai_rag_api.search_chunks(
p_rag_collection_id => 1,
p_query => 'How do I configure APEX authentication?',
p_limit => 10,
p_score_threshold => 0.7
));

Parameters:

ParameterTypeDefaultDescription
p_rag_collection_idNUMBERRequiredID of the RAG collection to search
p_queryVARCHAR2RequiredThe search query
p_limitNUMBERnullMaximum number of chunks to return. When null, the collection’s default_limit is used (falling back to 20).
p_score_thresholdNUMBERnullMinimum similarity score (0-1). When null, the collection’s default_score_threshold is used (falling back to 0.5).
p_rewrite_queryBOOLEANnullOverride: enable/disable query rewriting
p_rewrite_providerVARCHAR2nullOverride: provider for query rewriting
p_rewrite_modelVARCHAR2nullOverride: model for query rewriting
p_rewrite_configJSON_OBJECT_TnullOverride: generic UC AI config for query rewriting, merged over the collection’s query_rewriting.config (override keys win)

Use prepare_sources to get aggregated source documents. Matched chunks are grouped by document, documents are ranked by their best chunk score (so p_max_sources returns the most relevant ones), and each matched chunk is expanded by its surrounding chunks with overlapping ranges merged:

select document_id
, document_name
, version_number
, source_text
, chunk_count
, avg_score
from table(adm_ai_rag_api.prepare_sources(
p_rag_collection_id => 1,
p_query => 'How do I configure APEX authentication?',
p_limit => 10,
p_score_threshold => 0.7,
p_max_sources => 5
));

Additional Parameters:

ParameterTypeDefaultDescription
p_max_sourcesNUMBERnullMaximum number of source documents to return. When null, the collection’s default_max_sources is used (falling back to 5).
p_surrounding_chunksNUMBERnullNumber of neighboring chunks included around each matched chunk (0 disables expansion). When null, the collection’s default_surrounding_chunks is used (falling back to 1).

Use generate_answer to get an AI-generated response based on the matched sources. This requires the answer_model configuration in your collection:

declare
l_answer clob;
begin
l_answer := adm_ai_rag_api.generate_answer(
p_rag_collection_id => 1,
p_query => 'How do I configure APEX authentication?',
p_limit => 10,
p_score_threshold => 0.5,
p_max_sources => 3
);
dbms_output.put_line(l_answer);
end;

Additional Parameters:

ParameterTypeDefaultDescription
p_system_promptCLOBnullCustom system prompt for answer generation. When null, the collection’s answer_model.system_prompt is used (falling back to a built-in prompt).
p_answer_providerVARCHAR2nullOverride: provider for answer generation
p_answer_modelVARCHAR2nullOverride: model for answer generation
p_answer_configJSON_OBJECT_TnullOverride: generic UC AI config for answer generation, merged over the collection’s answer_model.config (override keys win)
p_rewrite_configJSON_OBJECT_TnullOverride: generic UC AI config for query rewriting, merged over the collection’s query_rewriting.config (override keys win)

For troubleshooting and analysis, use generate_answer_debug to also capture detailed debug information:

declare
l_answer clob;
l_debug_log_id number;
begin
adm_ai_rag_api.generate_answer_debug(
p_rag_collection_id => 1,
p_query => 'How do I configure APEX authentication?',
p_limit => 10,
p_score_threshold => 0.5,
p_max_sources => 3,
po_answer => l_answer,
po_debug_log_id => l_debug_log_id
);
dbms_output.put_line('Answer: ' || l_answer);
dbms_output.put_line('Debug Log ID: ' || l_debug_log_id);
end;

The debug log contains:

  • Original and rewritten queries
  • All matched chunks with scores
  • Prepared sources with aggregated text
  • System and user prompts sent to the AI
  • The AI response

Retrieve a formatted Markdown report of a debug log entry:

select adm_ai_rag_api.get_debug_log_markdown(p_debug_log_id => 123)
from dual;

You can customize the AI’s behavior by providing a custom system prompt:

declare
l_answer clob;
l_custom_prompt clob := 'You are a technical documentation expert. ' ||
'Provide concise, code-focused answers with examples. ' ||
'Always cite the source document name.';
begin
l_answer := adm_ai_rag_api.generate_answer(
p_rag_collection_id => 1,
p_query => 'How do I create a RESTful service?',
p_system_prompt => l_custom_prompt
);
dbms_output.put_line(l_answer);
end;

All query functions support runtime parameter overrides for query rewriting and answer generation, so you can use different models or credentials without changing the collection configuration:

declare
l_answer clob;
begin
l_answer := adm_ai_rag_api.generate_answer(
p_rag_collection_id => 1,
p_query => 'Explain the security model',
p_rewrite_query => true,
p_rewrite_provider => 'openai',
p_rewrite_model => 'gpt-4o-mini',
p_rewrite_config => json_object_t('{
"g_apex_web_credential": "OPENAI_API_KEY",
"g_base_url": "https://custom-rewrite-api.example.com/v1"
}'),
p_answer_provider => 'anthropic',
p_answer_model => 'claude-sonnet-4-5',
p_answer_config => json_object_t('{
"g_apex_web_credential": "ANTHROPIC_API_KEY",
"g_enable_reasoning": true,
"g_reasoning_level": "medium",
"g_base_url": "https://custom-api.example.com/v1"
}')
);
dbms_output.put_line(l_answer);
end;