Skip to main content

Proprietary framework

Enterprise AI QE Architecture

A structured architecture for testing and monitoring LLM, RAG, and agentic systems through evaluation datasets, prompt regression, retrieval testing, hallucination checks, and drift monitoring.

AI quality became evidence rather than opinion: one repeatable architecture for testing LLM, RAG and agentic systems before and after release.

Framework evidence

What the framework provides

  • AI-QE lifecycle
  • Evaluation scorecard (concept)
  • Regression suite design

CLIENT CONTEXT

GenAI systems routinely pass demos and fail in production, because they are not tested like enterprise software.

BUSINESS PROBLEM

LLM, RAG, and agentic systems need evaluation disciplines that traditional QE does not provide.

CONSTRAINTS

  • Probabilistic behaviour that traditional QE does not cover
  • Systems that pass a demonstration and fail in production
  • Evidence needed both before release and after it

ARCHITECTURE & APPROACH

Engineered VeriCore: evaluation datasets, prompt regression suites, retrieval testing, hallucination and grounding checks, and drift monitoring, pre-release and post-release.

WHAT CAPTIVOLT DELIVERED

Evaluation architecture · dataset design patterns · regression suite structure · scorecard model · monitoring approach.

ENGINEERING DECISIONS

  • Evaluation datasets versioned alongside the system they test
  • Prompt regression run as a suite rather than ad hoc
  • Retrieval quality measured separately from generation quality
  • Drift monitored after release, not only checked before it

EVIDENCE

Architecture walkthrough available.

In detail

How the evaluation architecture works.

The evaluation architecture

Six components. The ordering matters less than the separation: each answers a question the others cannot, and a missing one shows up as an evaluation that nobody trusts.

Evaluation datasets
Cases drawn from real business inputs, with the expected behaviour agreed by someone who owns the outcome, versioned alongside the system they test.
Fixtures and context snapshots
The retrieved context frozen with the case, so a failing run tells you whether the model or the retrieval changed.
Runners
Execution against a named build, model version and prompt version. A result that cannot name what produced it is an anecdote.
Scorers
One per dimension, separated deliberately: retrieval scored apart from generation, because a wrong answer from correct passages needs the opposite fix to a wrong answer from wrong ones.
Scorecard
The dimensions with their pass conditions and which are must-pass: the artefact a release decision is actually made against.
Monitors
The same dimensions sampled after release against the baseline the release produced, with breaches routed back into the dataset.

What a test case contains

The dataset is the part teams under-build. A case is not a prompt and an expected string.

Input
The question or task as a real user would put it, including the ones phrased badly, because those are where systems fail.
Context fixture
Which documents or records should be reachable, and under whose permissions, so an access failure is testable rather than incidental.
Expected behaviour
What a good answer must contain, must not contain, and must cite. Often "decline" is the correct behaviour and the case says so.
Provenance and owner
Where the case came from and who agreed it. An unowned case is deleted the first time it is inconvenient.
Version
Which release of the system it was agreed against, so a change in expectation is visible as a change rather than a mystery.

What is measured, and how

Twelve dimensions carry into VeriCore. These are the five whose method is most often misunderstood.

Groundedness
Each claim in the answer matched to a retrieved passage that supports it. Unmatched claims are counted, not averaged away.
Retrieval precision
Scored on the passages returned rather than on the answer written from them, so retrieval regressions are visible before they become answer regressions.
Task completion
For agents: did the workflow finish, on the cases agreed as in scope, with the side effects it should have had and none it should not.
Refusal correctness
Measured in both directions. Over-refusal is a failure too, and the one teams stop measuring once they have been embarrassed by an under-refusal.
Drift
The same suite re-run against the accepted baseline, which is why the baseline is an artefact rather than a memory.

What a regression run reports

The report design is the difference between a suite people read and a suite people mute.

Grouped by cause, not by case
One retrieval fault produces forty failing cases and one fix. A list of forty failures hides that; a list of causes shows it.
New failures separated from known ones
A run that is red for a reason already accepted is not the same as a run that broke something, and treating them alike is how suites get ignored.
The diff that produced it
Which prompt, model, tool or retrieval source changed since the last accepted run: the four things that change silently.
Verdict per must-pass dimension
Because that is what the gate reads, and a summary percentage is not.

The quality gate

A gate that has never returned a hold is a status meeting with a diagram.

Every must-pass dimension passes on this build
Not on a previous one, and not on a build with a different model version.
The adversarial suite has run against this version
Injection, leakage and refusal-bypass probes, on the version being released.
No new failure above the agreed severity
Known and accepted failures are named; anything new blocks by default.
Evidence attached to the release
The run, not a summary of the run in a message.
Release, hold, or a recorded override
An override is possible and is an exception with an accountable approver and an expiry date.

OUTCOME

A repeatable testing architecture that turns AI quality from opinion into evidence.

WHAT THE CLIENT OWNS NOW

  • The evaluation architecture and dataset design patterns
  • The regression suite structure
  • The scorecard model and monitoring approach

RELATED SOLUTION

Explore the capabilities behind the framework.

VeriCore · AI-Native SDLC Blueprint

← All real work

Want to see the artefacts?

Start with the work most relevant to your initiative.

Anonymised artefacts and reference discussions are available under NDA where client permission allows.