ACCELERATOR · Enterprise Retrieval & Grounding
Agentic RAG Accelerator
Permission-aware enterprise knowledge systems that connect LLMs to governed enterprise data, documents, systems, and workflows.
At a glance
The short answers.
- What is it?
- ACCELERATOR · Enterprise Retrieval & GroundingGround AI in controlled enterprise knowledge.
- Who is it for?
- Chief AI Officer / Head of AI
- What problem does it solve?
- Naive RAG leaks data, hallucinates, and degrades silently in production.
- What does it actually do?
- A reusable architecture plus delivery patterns covering the full knowledge pipeline, from ingestion to observed production behaviour.
- Where does it run?
- Your cloud, your identity provider, your network boundary. Retrieval runs under the caller’s permissions rather than a service account.
- What evidence does it produce?
- Citations against every answer, retrieval quality baselines, and a log of which sources were read under whose access.
- What does the client receive?
- Reference architecture and design records
- Working governed retrieval pipeline
- Evaluation baseline
- Operating runbooks
- How do we standardise RAG and agents?
One pipeline in a fixed order: thirteen steps from agreed sources to feedback, because almost every RAG failure is a step that was skipped.
The thirteen steps →- How do we manage evaluation?
As a stage of the pipeline itself: RAGAS-style evaluation, drift and cost monitoring, with the results fed back into the pipeline.
The RAG architecture →- How do we establish reusable AI capabilities?
By owning it: the pipeline, the configuration, the evaluation datasets and the baselines stay with you, and it runs, and changes, without us.
What you own afterwards →
Production RAG · ingestion → governance
A production RAG architecture.
How we design retrieval that stays governed, grounded, and observable in production, from ingestion through evaluation.
Ingest & govern
Parse and semantically chunk documents, then tag every chunk with permissions and lineage.
Retrieve with permissions
Hybrid dense + sparse retrieval that respects entitlements at the retrieval layer.
Rerank & ground
Cross-encoder reranking and context consolidation, with answers grounded in retrieved sources.
Generate with citations
Route to the right model, build the prompt, and return answers with source links.
Evaluate & observe
RAGAS-style evaluation, drift and cost monitoring, and feedback back into the pipeline.
Reference pipeline
Example output
Reference architecture.
Ingestion & indexing
Storage
Query & retrieval
Generation
Evaluation & governance
Illustrative reference architecture · representative stack, adapted per engagement · no client data shown.
The pipeline
Thirteen steps, and almost every RAG failure is one that was skipped.
Sources through feedback, in order. Most of these get mentioned on a vendor page; the order is the part that decides whether an answer can be trusted.
- 01
Sources
The systems and repositories we are allowed to read, named and agreed, not everything the crawler could reach.
- 02
Ingestion
Scheduled and event-driven pulls, with change detection so a re-run costs what changed rather than everything.
- 03
Parsing
Structure recovered before text: tables stay tables, headings stay headings, and a scanned page is OCR’d rather than skipped.
- 04
Chunking
Split on meaning rather than character count. A clause severed from its heading is a chunk that retrieves well and answers wrongly.
- 05
Metadata
Source, owner, effective date, supersession, classification and the permissions the original carries, attached at ingestion, not inferred later.
- 06
Embeddings & indexes
A vector index for meaning and a keyword index for exact terms, kept in step so both see the same corpus.
- 07
Permission-aware retrieval
The search runs under the caller’s identity. A document they may not read is not ranked, not counted and not summarised.
- 08
Reranking
A second pass over the shortlist that reads the question properly, because first-pass similarity is fast and frequently wrong at the top.
- 09
Context assembly
The passages that earn their place, within a budget, with their metadata intact so the model knows what it is holding.
- 10
LLM / agent
Answers from the passages supplied, and is expected to decline when they do not support one.
- 11
Citations
Each claim bound to the passage it came from, at generation time rather than matched back afterwards.
- 12
Evaluation
Groundedness, retrieval precision and answer quality, measured on real questions and run again whenever anything upstream changes.
- 13
Feedback
What people marked wrong, and what they searched for next, routed back into the evaluation set rather than into a backlog.
Video · 2 min
From scattered sources to a cited answer.
An illustrative contract question, followed through the retrieval system: sources connected and parsed, a contract chunked on its own clause boundaries, five retrieval modes reranked, a permission filter applied at retrieval so that two people asking get two different answers, a context package assembled for the model, and an answer that cites the page each statement came from.
Enterprise knowledge & RAG in action · 2:26
Read the video as text
- Knowledge everywhere: SharePoint documents, tickets and cases, CRM records, unstructured PDFs, email threads, knowledge-base articles and database tables. “Enterprise knowledge exists everywhere. Useful AI requires the right context — not simply more data.”
- Ingestion: connect the seven sources, extract text and tables, parse document structure, normalise to a common schema, classify by type and sensitivity, chunk into semantic units, enrich with metadata and entities, and index for keyword, vector and graph retrieval. “Connected. Understood. Structured. Ready for retrieval.”
- Boundary-aware chunking. A contract’s termination section cut at a fixed 512 tokens splits clause 7.2 mid-sentence and separates the early-termination fee table from its header. Chunked on its own boundaries, clauses 7.1, 7.2 and 7.3 stay whole, the confidentiality section starts a chunk of its own, and each chunk carries its metadata: section 7 of MSA-2024. “Preserve meaning before creating embeddings.”
- Multiple retrieval modes. The question: “What notice period applies if we terminate the Meridian contract early?” Keyword search, vector search, a metadata filter (contracts on the Meridian account), business facts (the start date and three-year term) and the relationship graph (the MSA and its Amendment 2) each contribute, and reranking by relevance, freshness and authority puts Amendment 2’s revised notice first, then the MSA’s notice period, then its fee table. “Retrieval is not one search.”
- Access-aware retrieval: the same question from two people. Legal counsel, with access to every contract document, retrieves the MSA clause and the legal-restricted Amendment 2, and the answer uses both: 60 days’ notice, as amended. A sales representative, with access to the account summary only, retrieves the MSA clause alone; the amendment is not retrieved, and the answer gives 90 days, the notice period in the documents the representative can read, and refers the question to Legal to confirm the current terms. “AI should never retrieve information the user is not authorised to see.”
- Context assembly: a context package for the contract notice question, built around the model from facts (term dates and parties), evidence (clause 7.1 and Amendment 2), relationships (the MSA to its amendment), history (no notice sent so far), policies (scoped to legal counsel) and signals (the renewal window is open). “Not a pile of documents. A purpose-built context package.”
- Grounded answer, with three citations. Either party may terminate for convenience with 90 days’ written notice [1]. Amendment 2 reduces notice to 60 days for renewals after 2026 [2]. An early termination fee of 15% applies in year 2 [3]. Each citation names its document, section and page in the legal library. “Answer → Evidence → Source.”
- “Enterprise RAG is Context Engineering.” Captivolt.
The parts that get conflated
Eight terms, and why each one exists.
Retrieval arguments usually turn out to be vocabulary arguments. These are what we mean, and what goes wrong when one is missing.
- Hybrid retrieval
Semantic and keyword search run together, and their results are merged before ranking.
Neither is sufficient. Vectors miss an exact part number; keywords miss a question asked in different words than the document uses.
- Exact reference
Lookup by identifier (contract number, SKU, policy reference, employee ID), routed to the keyword index or the system of record rather than to similarity.
An identifier has no semantic neighbourhood. Searching for one by meaning returns things that look like it, which is the worst possible answer.
- Semantic retrieval
Matching on meaning, so a question phrased in the business’s words finds a document written in the lawyer’s.
It is what makes retrieval useful on prose, and what makes it confidently wrong on identifiers. Hence both.
- Permissions
Retrieval inherits the permissions of the person asking, checked per request against the source system rather than against a copy.
A filter applied after ranking still leaks: through counts, through ordering, through a summary of documents the reader cannot open.
- Source attribution
Every claim carries the passage it came from, with document, section and version.
An answer nobody can check is unusable in a regulated decision, and indistinguishable from a fluent guess.
- Reranking
A slower, more accurate model re-scores the top candidates against the actual question before any of them reach the context.
First-pass retrieval optimises for speed across millions of chunks. The ordering it produces at the very top is the part most worth fixing.
- Evaluation
A suite over real questions with agreed answers, scoring retrieval and generation separately.
Separately matters: a wrong answer from correct passages and a wrong answer from wrong passages need opposite fixes.
- Monitoring
Groundedness sampling, retrieval quality against the release baseline, freshness of the corpus and what users did next.
Retrieval quality decays without any code changing: documents are added, superseded and re-permissioned underneath it.
Proof
One question, all the way through.
Including the two steps a retrieval demonstration never shows: the documents it was not allowed to read, and the sentence it declined to write.
“Which suppliers are on payment terms longer than 60 days, and what did we actually agree with Northwind?”
Deliberately a question that needs both halves: a filter over structured data, and a clause out of a contract nobody has read since it was signed.
- Routing
Split into an exact lookup over the vendor master for terms greater than 60 days, and a semantic search over contracts for Northwind’s payment clause.
- Permissions
Run under the requester, a procurement analyst. Two contracts matched and were excluded before ranking, as legal holds this analyst cannot open. They are not counted, mentioned or summarised in the answer.
- Retrieved evidence
Eleven vendor rows from the master; Northwind MSA §7.2 (payment terms) and Amendment 2 (which supersedes it).
- Reranking
The amendment outranks the original clause because it answers the question asked; the original is kept, marked superseded.
- Evaluation
Every sentence checked against a retrieved passage. One drafted sentence about early-settlement discount had no support and was removed rather than hedged.
- Response
What the analyst sees: the eleven suppliers, and Northwind’s terms as amended, with the amendment cited and the superseded clause noted. Nothing in it reveals the excluded contracts.
- Trace
The audit record, for authorised reviewers rather than the requester: question, identity, indexes searched, documents matched and excluded with the reason, passages used, the dropped sentence, and the corpus version, retained with the answer.
The reference flow worked through a concrete question, not a client’s query log, and with no confidence score, because a number we invented would be the least trustworthy thing on a page about traceable answers.
In practical terms
What gets installed, what you own, and how it connects.
The seven first questions are answered at the top of the page. These are the three a buyer asks next.
- What gets installed or configured?
- Retrieval and grounding services, the ingestion and chunking pipeline, an evaluation harness for retrieval quality, and connectors to the sources you approve.
- What do you own afterwards?
- The pipeline, the configuration, the evaluation datasets and the baselines. It runs, and changes, without us.
- How does it connect to Captivolt services?
- It is how the BUILD pillar delivers retrieval and grounding, rather than a separate purchase. Agentic AI & Engineering →
Modules
What is inside.
- Knowledge ingestion
- Permission-aware retrieval
- Source grounding
- Agent orchestration
- Response evaluation
- Observability
- Feedback loop
- Agentic RAG Framework: Reference architecture
Request an Agentic RAG Walkthrough.
We will walk through the architecture and how it maps onto your environment.