Proprietary framework
Enterprise AI QE Architecture
A structured architecture for testing and monitoring LLM, RAG, and agentic systems through evaluation datasets, prompt regression, retrieval testing, hallucination checks, and drift monitoring.
AI quality became evidence rather than opinion: one repeatable architecture for testing LLM, RAG and agentic systems before and after release.
What the framework provides
- AI-QE lifecycle
- Evaluation scorecard (concept)
- Regression suite design
CLIENT CONTEXT
GenAI systems routinely pass demos and fail in production, because they are not tested like enterprise software.
BUSINESS PROBLEM
LLM, RAG, and agentic systems need evaluation disciplines that traditional QE does not provide.
CONSTRAINTS
- Probabilistic behaviour that traditional QE does not cover
- Systems that pass a demonstration and fail in production
- Evidence needed both before release and after it
ARCHITECTURE & APPROACH
Engineered VeriCore: evaluation datasets, prompt regression suites, retrieval testing, hallucination and grounding checks, and drift monitoring, pre-release and post-release.
WHAT CAPTIVOLT DELIVERED
Evaluation architecture · dataset design patterns · regression suite structure · scorecard model · monitoring approach.
ENGINEERING DECISIONS
- Evaluation datasets versioned alongside the system they test
- Prompt regression run as a suite rather than ad hoc
- Retrieval quality measured separately from generation quality
- Drift monitored after release, not only checked before it
EVIDENCE
Architecture walkthrough available.
In detail
How the evaluation architecture works.
The evaluation architecture
Six components. The ordering matters less than the separation: each answers a question the others cannot, and a missing one shows up as an evaluation that nobody trusts.
- Evaluation datasets
- Cases drawn from real business inputs, with the expected behaviour agreed by someone who owns the outcome, versioned alongside the system they test.
- Fixtures and context snapshots
- The retrieved context frozen with the case, so a failing run tells you whether the model or the retrieval changed.
- Runners
- Execution against a named build, model version and prompt version. A result that cannot name what produced it is an anecdote.
- Scorers
- One per dimension, separated deliberately: retrieval scored apart from generation, because a wrong answer from correct passages needs the opposite fix to a wrong answer from wrong ones.
- Scorecard
- The dimensions with their pass conditions and which are must-pass: the artefact a release decision is actually made against.
- Monitors
- The same dimensions sampled after release against the baseline the release produced, with breaches routed back into the dataset.
What a test case contains
The dataset is the part teams under-build. A case is not a prompt and an expected string.
- Input
- The question or task as a real user would put it, including the ones phrased badly, because those are where systems fail.
- Context fixture
- Which documents or records should be reachable, and under whose permissions, so an access failure is testable rather than incidental.
- Expected behaviour
- What a good answer must contain, must not contain, and must cite. Often "decline" is the correct behaviour and the case says so.
- Provenance and owner
- Where the case came from and who agreed it. An unowned case is deleted the first time it is inconvenient.
- Version
- Which release of the system it was agreed against, so a change in expectation is visible as a change rather than a mystery.
What is measured, and how
Twelve dimensions carry into VeriCore. These are the five whose method is most often misunderstood.
- Groundedness
- Each claim in the answer matched to a retrieved passage that supports it. Unmatched claims are counted, not averaged away.
- Retrieval precision
- Scored on the passages returned rather than on the answer written from them, so retrieval regressions are visible before they become answer regressions.
- Task completion
- For agents: did the workflow finish, on the cases agreed as in scope, with the side effects it should have had and none it should not.
- Refusal correctness
- Measured in both directions. Over-refusal is a failure too, and the one teams stop measuring once they have been embarrassed by an under-refusal.
- Drift
- The same suite re-run against the accepted baseline, which is why the baseline is an artefact rather than a memory.
What a regression run reports
The report design is the difference between a suite people read and a suite people mute.
- Grouped by cause, not by case
- One retrieval fault produces forty failing cases and one fix. A list of forty failures hides that; a list of causes shows it.
- New failures separated from known ones
- A run that is red for a reason already accepted is not the same as a run that broke something, and treating them alike is how suites get ignored.
- The diff that produced it
- Which prompt, model, tool or retrieval source changed since the last accepted run: the four things that change silently.
- Verdict per must-pass dimension
- Because that is what the gate reads, and a summary percentage is not.
The quality gate
A gate that has never returned a hold is a status meeting with a diagram.
- Every must-pass dimension passes on this build
- Not on a previous one, and not on a build with a different model version.
- The adversarial suite has run against this version
- Injection, leakage and refusal-bypass probes, on the version being released.
- No new failure above the agreed severity
- Known and accepted failures are named; anything new blocks by default.
- Evidence attached to the release
- The run, not a summary of the run in a message.
- Release, hold, or a recorded override
- An override is possible and is an exception with an accountable approver and an expiry date.
OUTCOME
A repeatable testing architecture that turns AI quality from opinion into evidence.
WHAT THE CLIENT OWNS NOW
- The evaluation architecture and dataset design patterns
- The regression suite structure
- The scorecard model and monitoring approach
RELATED SOLUTION
Explore the capabilities behind the framework.
VeriCore · AI-Native SDLC Blueprint
Want to see the artefacts?
Start with the work most relevant to your initiative.
Anonymised artefacts and reference discussions are available under NDA where client permission allows.