Skip to main content

ACCELERATOR · Enterprise Retrieval & Grounding

Agentic RAG Accelerator

Permission-aware enterprise knowledge systems that connect LLMs to governed enterprise data, documents, systems, and workflows.

At a glance

The short answers.

What is it?
ACCELERATOR · Enterprise Retrieval & GroundingGround AI in controlled enterprise knowledge.
Who is it for?
Chief AI Officer / Head of AI
What problem does it solve?
Naive RAG leaks data, hallucinates, and degrades silently in production.
What does it actually do?
A reusable architecture plus delivery patterns covering the full knowledge pipeline, from ingestion to observed production behaviour.
Where does it run?
Your cloud, your identity provider, your network boundary. Retrieval runs under the caller’s permissions rather than a service account.
What evidence does it produce?
Citations against every answer, retrieval quality baselines, and a log of which sources were read under whose access.
What does the client receive?
  • Reference architecture and design records
  • Working governed retrieval pipeline
  • Evaluation baseline
  • Operating runbooks
How do we standardise RAG and agents?

One pipeline in a fixed order: thirteen steps from agreed sources to feedback, because almost every RAG failure is a step that was skipped.

The thirteen steps →
How do we manage evaluation?

As a stage of the pipeline itself: RAGAS-style evaluation, drift and cost monitoring, with the results fed back into the pipeline.

The RAG architecture →
How do we establish reusable AI capabilities?

By owning it: the pipeline, the configuration, the evaluation datasets and the baselines stay with you, and it runs, and changes, without us.

What you own afterwards →

Production RAG · ingestion → governance

A production RAG architecture.

How we design retrieval that stays governed, grounded, and observable in production, from ingestion through evaluation.

  1. Ingest & govern

    Parse and semantically chunk documents, then tag every chunk with permissions and lineage.

  2. Retrieve with permissions

    Hybrid dense + sparse retrieval that respects entitlements at the retrieval layer.

  3. Rerank & ground

    Cross-encoder reranking and context consolidation, with answers grounded in retrieved sources.

  4. Generate with citations

    Route to the right model, build the prompt, and return answers with source links.

  5. Evaluate & observe

    RAGAS-style evaluation, drift and cost monitoring, and feedback back into the pipeline.

Reference pipeline

Governed data layersBronze · Silver · Gold
Ingestion & semantic chunkingstructure-aware parsing
Embedding & vector indexhybrid dense + sparse
Permission-aware retrievalRBAC at the retrieval layer
Rerank & groundcross-encoder · source grounding
Grounded responseanswer + citations
GUARDRAILS · ACCESS CONTROL · EVALUATION · OBSERVABILITY · FEEDBACK LOOP

Example output

Reference architecture.

REFERENCE ARCHITECTUREProduction RAG · ingestion → governance

Ingestion & indexing

SourcesS3 · API · files
Extract & chunksemantic chunking
Metadata + RBACpermissions · lineage
Embed & indexdense vectors

Storage

Vector storeQdrant / PGVector
Sparse indexBM25 / Elastic
Document storeobject + records

Query & retrieval

Query routingintent-aware
Hybrid retrievaldense + sparse
Cross-encoder rerankrelevance
Human-in-loop gatehigh-risk review

Generation

Model routercost-tiered
Prompt buildercontext + query
LLM + fallbackprimary + backup
Citation buildersource links

Evaluation & governance

Faithfulness eval
Lineage & audit
Cost monitoring
PII / redaction

Illustrative reference architecture · representative stack, adapted per engagement · no client data shown.

The pipeline

Thirteen steps, and almost every RAG failure is one that was skipped.

Sources through feedback, in order. Most of these get mentioned on a vendor page; the order is the part that decides whether an answer can be trusted.

  1. 01

    Sources

    The systems and repositories we are allowed to read, named and agreed, not everything the crawler could reach.

  2. 02

    Ingestion

    Scheduled and event-driven pulls, with change detection so a re-run costs what changed rather than everything.

  3. 03

    Parsing

    Structure recovered before text: tables stay tables, headings stay headings, and a scanned page is OCR’d rather than skipped.

  4. 04

    Chunking

    Split on meaning rather than character count. A clause severed from its heading is a chunk that retrieves well and answers wrongly.

  5. 05

    Metadata

    Source, owner, effective date, supersession, classification and the permissions the original carries, attached at ingestion, not inferred later.

  6. 06

    Embeddings & indexes

    A vector index for meaning and a keyword index for exact terms, kept in step so both see the same corpus.

  7. 07

    Permission-aware retrieval

    The search runs under the caller’s identity. A document they may not read is not ranked, not counted and not summarised.

  8. 08

    Reranking

    A second pass over the shortlist that reads the question properly, because first-pass similarity is fast and frequently wrong at the top.

  9. 09

    Context assembly

    The passages that earn their place, within a budget, with their metadata intact so the model knows what it is holding.

  10. 10

    LLM / agent

    Answers from the passages supplied, and is expected to decline when they do not support one.

  11. 11

    Citations

    Each claim bound to the passage it came from, at generation time rather than matched back afterwards.

  12. 12

    Evaluation

    Groundedness, retrieval precision and answer quality, measured on real questions and run again whenever anything upstream changes.

  13. 13

    Feedback

    What people marked wrong, and what they searched for next, routed back into the evaluation set rather than into a backlog.

Video · 2 min

From scattered sources to a cited answer.

An illustrative contract question, followed through the retrieval system: sources connected and parsed, a contract chunked on its own clause boundaries, five retrieval modes reranked, a permission filter applied at retrieval so that two people asking get two different answers, a context package assembled for the model, and an answer that cites the page each statement came from.

Enterprise knowledge & RAG in action · 2:26

Read the video as text
  1. Knowledge everywhere: SharePoint documents, tickets and cases, CRM records, unstructured PDFs, email threads, knowledge-base articles and database tables. “Enterprise knowledge exists everywhere. Useful AI requires the right context — not simply more data.”
  2. Ingestion: connect the seven sources, extract text and tables, parse document structure, normalise to a common schema, classify by type and sensitivity, chunk into semantic units, enrich with metadata and entities, and index for keyword, vector and graph retrieval. “Connected. Understood. Structured. Ready for retrieval.”
  3. Boundary-aware chunking. A contract’s termination section cut at a fixed 512 tokens splits clause 7.2 mid-sentence and separates the early-termination fee table from its header. Chunked on its own boundaries, clauses 7.1, 7.2 and 7.3 stay whole, the confidentiality section starts a chunk of its own, and each chunk carries its metadata: section 7 of MSA-2024. “Preserve meaning before creating embeddings.”
  4. Multiple retrieval modes. The question: “What notice period applies if we terminate the Meridian contract early?” Keyword search, vector search, a metadata filter (contracts on the Meridian account), business facts (the start date and three-year term) and the relationship graph (the MSA and its Amendment 2) each contribute, and reranking by relevance, freshness and authority puts Amendment 2’s revised notice first, then the MSA’s notice period, then its fee table. “Retrieval is not one search.”
  5. Access-aware retrieval: the same question from two people. Legal counsel, with access to every contract document, retrieves the MSA clause and the legal-restricted Amendment 2, and the answer uses both: 60 days’ notice, as amended. A sales representative, with access to the account summary only, retrieves the MSA clause alone; the amendment is not retrieved, and the answer gives 90 days, the notice period in the documents the representative can read, and refers the question to Legal to confirm the current terms. “AI should never retrieve information the user is not authorised to see.”
  6. Context assembly: a context package for the contract notice question, built around the model from facts (term dates and parties), evidence (clause 7.1 and Amendment 2), relationships (the MSA to its amendment), history (no notice sent so far), policies (scoped to legal counsel) and signals (the renewal window is open). “Not a pile of documents. A purpose-built context package.”
  7. Grounded answer, with three citations. Either party may terminate for convenience with 90 days’ written notice [1]. Amendment 2 reduces notice to 60 days for renewals after 2026 [2]. An early termination fee of 15% applies in year 2 [3]. Each citation names its document, section and page in the legal library. “Answer → Evidence → Source.”
  8. “Enterprise RAG is Context Engineering.” Captivolt.

The parts that get conflated

Eight terms, and why each one exists.

Retrieval arguments usually turn out to be vocabulary arguments. These are what we mean, and what goes wrong when one is missing.

Hybrid retrieval

Semantic and keyword search run together, and their results are merged before ranking.

Neither is sufficient. Vectors miss an exact part number; keywords miss a question asked in different words than the document uses.

Exact reference

Lookup by identifier (contract number, SKU, policy reference, employee ID), routed to the keyword index or the system of record rather than to similarity.

An identifier has no semantic neighbourhood. Searching for one by meaning returns things that look like it, which is the worst possible answer.

Semantic retrieval

Matching on meaning, so a question phrased in the business’s words finds a document written in the lawyer’s.

It is what makes retrieval useful on prose, and what makes it confidently wrong on identifiers. Hence both.

Permissions

Retrieval inherits the permissions of the person asking, checked per request against the source system rather than against a copy.

A filter applied after ranking still leaks: through counts, through ordering, through a summary of documents the reader cannot open.

Source attribution

Every claim carries the passage it came from, with document, section and version.

An answer nobody can check is unusable in a regulated decision, and indistinguishable from a fluent guess.

Reranking

A slower, more accurate model re-scores the top candidates against the actual question before any of them reach the context.

First-pass retrieval optimises for speed across millions of chunks. The ordering it produces at the very top is the part most worth fixing.

Evaluation

A suite over real questions with agreed answers, scoring retrieval and generation separately.

Separately matters: a wrong answer from correct passages and a wrong answer from wrong passages need opposite fixes.

Monitoring

Groundedness sampling, retrieval quality against the release baseline, freshness of the corpus and what users did next.

Retrieval quality decays without any code changing: documents are added, superseded and re-permissioned underneath it.

Proof

One question, all the way through.

Including the two steps a retrieval demonstration never shows: the documents it was not allowed to read, and the sentence it declined to write.

The question

“Which suppliers are on payment terms longer than 60 days, and what did we actually agree with Northwind?”

Deliberately a question that needs both halves: a filter over structured data, and a clause out of a contract nobody has read since it was signed.

  1. Routing

    Split into an exact lookup over the vendor master for terms greater than 60 days, and a semantic search over contracts for Northwind’s payment clause.

  2. Permissions

    Run under the requester, a procurement analyst. Two contracts matched and were excluded before ranking, as legal holds this analyst cannot open. They are not counted, mentioned or summarised in the answer.

  3. Retrieved evidence

    Eleven vendor rows from the master; Northwind MSA §7.2 (payment terms) and Amendment 2 (which supersedes it).

  4. Reranking

    The amendment outranks the original clause because it answers the question asked; the original is kept, marked superseded.

  5. Evaluation

    Every sentence checked against a retrieved passage. One drafted sentence about early-settlement discount had no support and was removed rather than hedged.

  6. Response

    What the analyst sees: the eleven suppliers, and Northwind’s terms as amended, with the amendment cited and the superseded clause noted. Nothing in it reveals the excluded contracts.

  7. Trace

    The audit record, for authorised reviewers rather than the requester: question, identity, indexes searched, documents matched and excluded with the reason, passages used, the dropped sentence, and the corpus version, retained with the answer.

The reference flow worked through a concrete question, not a client’s query log, and with no confidence score, because a number we invented would be the least trustworthy thing on a page about traceable answers.

In practical terms

What gets installed, what you own, and how it connects.

The seven first questions are answered at the top of the page. These are the three a buyer asks next.

What gets installed or configured?
Retrieval and grounding services, the ingestion and chunking pipeline, an evaluation harness for retrieval quality, and connectors to the sources you approve.
What do you own afterwards?
The pipeline, the configuration, the evaluation datasets and the baselines. It runs, and changes, without us.
How does it connect to Captivolt services?
It is how the BUILD pillar delivers retrieval and grounding, rather than a separate purchase. Agentic AI & Engineering →

Modules

What is inside.

  • Knowledge ingestion
  • Permission-aware retrieval
  • Source grounding
  • Agent orchestration
  • Response evaluation
  • Observability
  • Feedback loop
REAL WORK
RELATED · BUILD · Enterprise AI Agents · VeriCore

Request an Agentic RAG Walkthrough.

We will walk through the architecture and how it maps onto your environment.