Enterprise AI QE Architecture / VeriCore
Establish the testing architecture for LLM, RAG, and agentic systems across development and production.
ASSURE
Captivolt helps enterprises validate AI behaviour, govern risk, apply security controls to agentic systems, and create the evidence required for responsible deployment.
What data classification and access policy allow. Retrieval is permission-aware, with redaction and filtering, and it is tested for data leakage.
AI security →At a release gate that can say no: every must-pass dimension has to pass on the build being released, with no open finding above the agreed severity.
The release gate →By the controls as they run, not for a meeting afterwards: a scorecard, a gate decision, red-team results, a governance record and a monitoring result.
The five artefacts →By watching the same dimensions in production that were measured before release, and re-evaluating on drift, on any change, and on a schedule.
The assurance lifecycle →The business problem
AI is entering production faster than the quality, governance, and security disciplines needed to oversee it, and regulators, boards, and auditors are starting to ask for evidence.
Why current approaches fail
Four habits that let a system reach production on confidence rather than evidence.
A convincing demonstration runs on curated inputs and friendly users. Production has neither.
A suite that last ran against a previous build says nothing about the version being released.
A policy with no register, owner or evidence behind it cannot answer when an auditor asks what is running.
Behaviour drifts as data, prompts and models change, and without an accepted baseline nobody sees it happen.
Captivolt point of view
A gate that cannot hold a release is a status meeting with a diagram.
Evidence assembled for the meeting evidences nothing an auditor can rely on.
It shows how a system behaved on the cases that were run. Accepting what remains stays with an accountable person.
The assurance lifecycle
AI should move to production because its behaviour has been measured and its risks are controlled, not because the demonstration looked convincing.
Agree what good means for this system, and what would stop a release, before it is built, while the answer can still be honest.
Evaluation criteria and risk classification
Measure behaviour against those criteria, on real cases from the business rather than examples chosen to pass.
Evaluation scorecard
Run the adversarial and regression suites against this version. A suite that last ran against a previous build has told you nothing about this one.
Red-team and regression results
Compare the results to the release conditions. The gate has to be able to say no, or it is a status meeting with a diagram.
Release gate decision
A named, accountable person accepts the residual risk against the evidence. Not the delivery team, and not a committee in the abstract.
Governance record entry
The controls go live with the system rather than after it: permissions, filters, limits and logging are part of the release.
Controls running in the environment
Watch the same dimensions in production that were measured before it. Different measures before and after release means neither can be compared.
Monitoring result
On drift, on any change to model, prompt, tools or data, and on a schedule regardless. A passing score has a date on it.
A new accepted baseline, or a hold
01 Define · 02 Evaluate · 03 Test · 04 Gate · 07 Monitor · 08 Re-evaluate
The measurement side: evaluation harness and datasets, regression baselines, scorecards, release gates and drift monitoring.
Explore VeriCore →01 Define · 05 Approve · 07 Monitor
The record side: system register, risk classification, intake and approval workflow, and the evidence model underneath it.
Explore AegisIQ →06 Deploy
Deploy is neither accelerator. Controls have to be built into the running system, which is engineering work, so it is named here and owned there.
How controls get built in →Video · 3 min
Guardrails, evaluations, observability and compliance working on one AI application in production, with human oversight around all four: a request stopped at the data boundary, a groundedness regression caught between releases, a trace with its latency and cost, and the governance record the interaction leaves. The scores on screen are illustrative, not a client’s.
How we make AI trustworthy · 2:44
Testing architecture
AI-QE architecture
Capabilities
Establish the testing architecture for LLM, RAG, and agentic systems across development and production.
Validate correctness, groundedness, consistency, relevance, safety, and policy alignment.
Test retrieval precision, source grounding, access control, hallucination risk, and response reliability.
Evaluate task completion, tool selection, tool input quality, instruction adherence, and workflow reliability.
Detect failures when prompts, models, context, tools, or retrieval sources change.
Test for jailbreaks, prompt injection, unsafe outputs, data leakage, and adversarial behaviour.
Build AI inventory, risk classification, policy controls, approval workflows, and evidence models.
Create a living register of AI systems, owners, models, data sources, integrations, and risk levels.
Assess privacy, fairness, safety, explainability, business impact, and operational risk.
Define policies, review boards, evidence requirements, decision rights, and accountability structures.
Govern retrieval sources, access permissions, lineage, grounding quality, and leakage risk.
Prepare AI management-system controls, records, policies, and operating evidence.
Map AI systems and data practices to relevant regulatory expectations.
Detect and reduce prompt injection, jailbreak, tool-abuse and unsafe-action risk in agents and LLM applications.
Reduce unauthorised exposure through permission-aware retrieval, redaction, filtering, and testing.
Align data classification, access policies, and model usage controls.
Define playbooks, escalation paths, and recovery processes for AI-related incidents.
Assess code, APIs, architecture, and cloud environments for exploitable risk.
Identify design-level risks before systems are built or released.
Monitor quality, behaviour, drift, safety, cost, and performance after deployment.
Track operational cost, latency, token usage, and reliability against agreed thresholds.
Route production findings back into evaluation suites, prompts, and retrieval improvements.
Build security governance, risk treatment, Statement of Applicability, Annex A controls, and audit evidence.
Assess vendors, platforms, and technology dependencies for security and compliance risk.
Provide senior security leadership for risk, control maturity, board reporting, and audit readiness.
Prepare control evidence and operating cadence for SOC 2 readiness.
Reduce duplicated compliance work by mapping controls across ISO, SOC 2, DPDP, GDPR, and sector expectations.
Use cases
How Captivolt delivers
Typically a governance or AI-QE gap assessment, followed by framework implementation and an ongoing assurance cadence.
The evidence
Five artefacts, laid out as the records they are. Four of them are structures rather than filled-in examples, because the values belong to your system, and a page arguing that untraceable claims are worthless has no business printing invented ones.
Every dimension carries a pass condition, and most of them are must-pass: a release cannot proceed while one fails.
The dimensions and pass conditions are ours; the scores are your system’s, measured on your cases. We do not publish example numbers, because a number nobody can trace is the thing this page argues against.
The conditions checked before a version ships, and the three answers the gate is allowed to give.
A gate that can only say yes is a status meeting. These are the conditions; the outcome is whatever the evidence produces on the day.
Six probes we run against agentic and RAG systems, each paired with the control that is supposed to stop it.
Indirect prompt injection through retrieved content
A document in the corpus carries instructions and the agent follows them, having treated retrieved text as direction rather than as data.
Retrieved content is never executed as instruction, and a tool call is authorised by policy rather than by the model’s intent.
Permission escalation through retrieval
One user surfaces content only another is entitled to see, because the index was searched with the service’s access rather than the caller’s.
Retrieval inherits the caller’s permissions and is checked per request, not per index build.
Tool use beyond remit
A tool outside the task is called, or called with inputs nobody supplied, or called more times than the work requires.
An allow-list registry, schema validation on every input, and an authority boundary the model has no way to widen.
Exfiltration through the response or an outbound call
Credentials, secrets or personal data reach the output, a log, or a request to an external system.
Output filtering, an egress policy on every tool, and redaction at the retrieval boundary.
Refusal bypass by reframing
A restricted request succeeds once it is rephrased: as fiction, as translation, as code, as a hypothetical.
Policy evaluated outside the prompt, where a reframing has nothing to argue with.
Confident answers from superseded context
A withdrawn policy or an expired price is quoted as current, with a citation that makes it look verified.
Freshness and supersession checks on the corpus, and no answer the retrieved evidence does not support.
Public attack classes, named in full and described at the level of the class and its defence. Findings from an engagement belong to the client and stay there.
What is held for each AI system, so an auditor’s question has an answer that already exists.
The record structure, as implemented in AegisIQ. Populated from your systems: a register we filled in with examples would be a worked fiction.
The signals watched after release, and the escalation path when one of them moves.
Thresholds are set per system at Define, against the baseline that release produced. That is why they are not printed here as figures.
Evidence
Relevant accelerators
It is the measurement side of the ASSURE lifecycle: stages 02 Evaluate through 04 Gate, and 07 Monitor.
It is the record side of the ASSURE lifecycle: stage 05 Approve, and the governance record every other stage writes to.
A structured first conversation about what you are trying to build, govern, or scale.