Skip to main content

Agentic AI & RAG

Why enterprise RAG fails in production

Most RAG systems fail because they treat retrieval as a search problem instead of an enterprise architecture problem.

CAPTIVOLT INSIGHTS · Published · Updated · 5 min read

Executive summary

Most RAG systems fail because they treat retrieval as a search problem instead of an enterprise architecture problem. Production RAG requires permission-aware retrieval, source grounding, quality evaluation, metadata discipline, and operating feedback loops. Teams that engineer these from the start ship systems that hold; teams that bolt them on later spend quarters firefighting.

The problem

The demo works. Users ask questions, answers come back, leadership approves the rollout. Then production arrives: a finance document surfaces to someone outside finance; answers cite documents superseded months ago; quality drifts as the corpus grows and nobody notices until users stop trusting the system. None of these failures appears in a demo, because demos run on curated data, with friendly users, and without time.

A practical framework

  1. 01

    Treat retrieval as architecture, not search. Ingestion, indexing, chunking, and metadata are engineering decisions with failure modes: design them deliberately.

  2. 02

    Make retrieval permission-aware at the retrieval layer. Access control applied after retrieval has already leaked; the index itself must respect entitlements.

  3. 03

    Ground every answer in traceable sources. If the system cannot show where an answer came from, users cannot trust it and auditors cannot accept it.

  4. 04

    Build an evaluation baseline before launch. Retrieval precision, groundedness, and hallucination rates need a measured starting point, or drift is invisible.

  5. 05

    Close the loop in operation. Production findings must flow back into evaluation suites, metadata fixes, and retrieval improvements. Silently decaying systems are the norm, not the exception.

Going further

Ingestion & parsing

A taxonomy of enterprise RAG failure, organised by the stage where it breaks, because the fix lives at that stage, and most failures are diagnosed three stages downstream of their cause. These first ones happen before a question is asked, so no demonstration ever shows them.

Silent parse loss

Tables flattened into run-on text, scanned pages skipped, headers and footers ingested as body. The corpus looks complete and answers from pages it cannot actually read.

Detect: sample parsed output against the source by document type, and count pages with no extractable text.

Stale corpus

No change detection, so re-ingestion is skipped or becomes a full rebuild nobody schedules. The index describes last quarter.

Detect: compare source modification dates against index timestamps, and alert on documents changed but not re-indexed.

Chunking & metadata

Where meaning is destroyed without anything visibly breaking.

Boundary severance

A clause split from the condition that governs it, a table row from its header. Each chunk embeds well and answers wrongly.

Detect: boundary-spanning questions in the evaluation set: ones whose answer needs text from both sides of a split.

Missing supersession

No effective date or superseded-by metadata, so a withdrawn policy retrieves as confidently as its replacement.

Detect: seed the evaluation set with questions whose correct answer has changed, and check which version is cited.

Lost permissions metadata

The source document’s entitlements are not carried onto its chunks, so the index cannot filter by who is asking even when retrieval tries to.

Detect: for any chunk, try to reconstruct the original document’s access rule. If it cannot be done, filtering is guesswork.

Retrieval & permissions

The stage most teams think RAG is, and where the most expensive failures live.

Similarity search on an identifier

A contract number, SKU or policy reference sent to vector search, which returns things that look like it.

Detect: exact-reference questions in the evaluation set. A semantic near-miss is a failure, not a partial score.

Retrieval under a service account

The search runs with more rights than the person asking. Nothing leaks until somebody phrases the right question.

Detect: run the evaluation set as users with different entitlements and compare what each can surface.

Lost in the middle

The right passage is retrieved but ranked where the model attends least, so it is present and ignored.

Detect: measure answer accuracy against the position of the supporting passage in the assembled context.

Assembly & generation

Retrieval can be right and the answer still wrong.

Context over-stuffing

More passages sent than the question needs, so the relevant one competes with nine that are merely similar.

Detect: plot answer quality against the number of passages supplied. If it falls as context grows, the budget is wrong.

Unresolved conflict

Two retrieved passages disagree and the model blends them into an answer neither of them supports.

Detect: evaluation cases built from known contradictions, checking that the answer names the conflict rather than averaging it.

Unsupported synthesis

The answer goes beyond the evidence: fluent, plausible, and drawn from the model’s general knowledge rather than the corpus.

Detect: claim-level groundedness. Every sentence matched to a supporting passage, and unmatched sentences counted rather than averaged away.

Citation & operations

Failures that make a correct answer unusable, or a working system quietly worse.

Post-hoc citation

Citations matched back to the answer after generation, so a claim can carry a source that does not actually say it.

Detect: verify that each cited passage supports its claim. A citation that does not support its sentence is worse than none.

Silent decay

No accepted baseline, so retrieval quality falls as documents are added, superseded and re-permissioned, and nobody sees the curve.

Detect: re-run the same suite on a schedule against a stored baseline. Decay is only visible against a fixed reference.

Practical implications

  • RAG failures are architecture failures, not model failures.
  • Permissions belong inside the retrieval layer.
  • No evaluation baseline means no drift detection.
  • Plan the operating feedback loop before go-live, not after the first incident.

About this article

Author
Captivolt Insights
Published
· updated

References

  1. Lost in the Middle: How Language Models Use Long Contexts · Liu et al., arXiv:2307.03172