Skip to main content

ASSURE

A convincing demo is not evidence.

Captivolt helps enterprises validate AI behaviour, govern risk, apply security controls to agentic systems, and create the evidence required for responsible deployment.

The short answers.

What data can AI access?

What data classification and access policy allow. Retrieval is permission-aware, with redaction and filtering, and it is tested for data leakage.

AI security →
How are controls enforced?

At a release gate that can say no: every must-pass dimension has to pass on the build being released, with no open finding above the agreed severity.

The release gate →
How is evidence created?

By the controls as they run, not for a meeting afterwards: a scorecard, a gate decision, red-team results, a governance record and a monitoring result.

The five artefacts →
How do we detect failures?

By watching the same dimensions in production that were measured before release, and re-evaluating on drift, on any change, and on a schedule.

The assurance lifecycle →

The business problem

AI Quality, Governance & Security

AI is entering production faster than the quality, governance, and security disciplines needed to oversee it, and regulators, boards, and auditors are starting to ask for evidence.

Why current approaches fail

AI quality is usually judged, not measured.

Four habits that let a system reach production on confidence rather than evidence.

Judged in a demonstration

A convincing demonstration runs on curated inputs and friendly users. Production has neither.

Tested once

A suite that last ran against a previous build says nothing about the version being released.

Governed on paper

A policy with no register, owner or evidence behind it cannot answer when an auditor asks what is running.

Unwatched after release

Behaviour drifts as data, prompts and models change, and without an accepted baseline nobody sees it happen.

Captivolt point of view

AI quality is an engineering discipline, not an opinion formed in a demo. Probabilistic systems need evaluation datasets, regression suites and drift monitoring the way software needed unit tests.

A release gate has to be able to say no.

A gate that cannot hold a release is a status meeting with a diagram.

Evidence is produced by the controls as they run.

Evidence assembled for the meeting evidences nothing an auditor can rely on.

Measurement is not a promise of safety.

It shows how a system behaved on the cases that were run. Accepting what remains stays with an accountable person.

The assurance lifecycle

Eight stages, and the last one goes back to the second.

AI should move to production because its behaviour has been measured and its risks are controlled, not because the demonstration looked convincing.

  1. 01

    Define

    Agree what good means for this system, and what would stop a release, before it is built, while the answer can still be honest.

    Produces

    Evaluation criteria and risk classification

    VeriCoreAegisIQ
  2. 02

    Evaluate

    Measure behaviour against those criteria, on real cases from the business rather than examples chosen to pass.

    Produces

    Evaluation scorecard

    VeriCore
  3. 03

    Test

    Run the adversarial and regression suites against this version. A suite that last ran against a previous build has told you nothing about this one.

    Produces

    Red-team and regression results

    VeriCore
  4. 04

    Gate

    Compare the results to the release conditions. The gate has to be able to say no, or it is a status meeting with a diagram.

    Produces

    Release gate decision

    VeriCore
  5. 05

    Approve

    A named, accountable person accepts the residual risk against the evidence. Not the delivery team, and not a committee in the abstract.

    Produces

    Governance record entry

    AegisIQ
  6. 06

    Deploy

    The controls go live with the system rather than after it: permissions, filters, limits and logging are part of the release.

    Produces

    Controls running in the environment

    Delivery
  7. 07

    Monitor

    Watch the same dimensions in production that were measured before it. Different measures before and after release means neither can be compared.

    Produces

    Monitoring result

    VeriCoreAegisIQ
  8. 08

    Re-evaluate

    On drift, on any change to model, prompt, tools or data, and on a schedule regardless. A passing score has a date on it.

    Produces

    A new accepted baseline, or a hold

    VeriCore
    ↩ returns to 02 Evaluate

VeriCore

01 Define · 02 Evaluate · 03 Test · 04 Gate · 07 Monitor · 08 Re-evaluate

The measurement side: evaluation harness and datasets, regression baselines, scorecards, release gates and drift monitoring.

Explore VeriCore →

AegisIQ

01 Define · 05 Approve · 07 Monitor

The record side: system register, risk classification, intake and approval workflow, and the evidence model underneath it.

Explore AegisIQ →

Delivery

06 Deploy

Deploy is neither accelerator. Controls have to be built into the running system, which is engineering work, so it is named here and owned there.

How controls get built in →

Video · 3 min

Four controls, one assurance layer.

Guardrails, evaluations, observability and compliance working on one AI application in production, with human oversight around all four: a request stopped at the data boundary, a groundedness regression caught between releases, a trace with its latency and cost, and the governance record the interaction leaves. The scores on screen are illustrative, not a client’s.

How we make AI trustworthy · 2:44

Read the video as text
  1. The question. Four risks around an AI application in production: can it access something it shouldn’t? Is the answer actually grounded? Has its behaviour changed? Can we prove what happened? “Accuracy alone is not enough.”
  2. Trust is a system: guardrails (boundaries), evals (quality), observability (visibility) and compliance (evidence) around the application, with human oversight around all four. “Trust is engineered — not declared.”
  3. Guardrails. A sensitive request (“Export this customer’s contract terms and email them to our external partner”) passes identity (an account manager, verified through SSO), data permission (the contract terms are within the user’s scope) and prompt policy (no prohibited content or injection), and is stopped at tool permission: external email is not permitted for this data class. The action is blocked, nothing is sent, and the event is logged. “Control what AI can access, generate and execute.”
  4. Evals. An evaluation pipeline scores each response before deployment and continuously after it. In the illustration, completeness, citation accuracy, relevance, policy adherence and task success pass, while groundedness falls below its 0.85 threshold at release v2.3 and is flagged as a regression. “Measure whether the AI is actually working.”
  5. Observability. One request traced from the user through the agent, retrieval, the model and a tool to the response, with the time each step took; the tool call fails and is retried. The trace keeps the latency, the model used, the tokens, the cost, the sources retrieved and the errors. “Understand what happened inside every interaction.”
  6. Compliance. A governance record of the interaction: the user (an account manager, scoped), the purpose (expansion planning), the data accessed (CRM, the MSA and a usage report), the policy applied (data access and an external-send block), the model used, the approval (not required to read or recommend), the evidence (the trace and the evaluation run) and the time. Traceable for security, risk, compliance and audit. “Turn AI behaviour into evidence.”
  7. Human control. Reading is low risk and automated; recommending is medium risk and automated; executing is high risk and needs approval, routed to an approver and logged. “Human control must be operational — not aspirational.”
  8. One assurance layer (guardrails, evals, observability and compliance) around the AI application, under human oversight. “Production AI you can operate — and defend.” Captivolt.

Testing architecture

What runs inside Evaluate and Test.

The scorecard says what is measured and what passing means; this is the machinery that produces it: datasets versioned with the system, suites rather than manual passes, and a gate that reads a score instead of a room.

AI-QE architecture

Evaluation datasetsreal cases, versioned with the system
Prompt & retrieval regressionrun as a suite, not by hand
Agent & workflow evalstask completion · tool-use accuracy
Red teaming & safetyprompt injection · jailbreak · leakage
Release gatescored against a threshold, not an opinion
Monitoring after releasedrift · cost · latency · feedback loops
AI SYSTEM INVENTORY · RISK CLASSIFICATION · AUDIT EVIDENCE · INCIDENT RESPONSE

Capabilities

What this pillar covers.

AI Quality & Evaluation

Enterprise AI QE Architecture / VeriCore

Establish the testing architecture for LLM, RAG, and agentic systems across development and production.

LLM & GenAI Validation

Validate correctness, groundedness, consistency, relevance, safety, and policy alignment.

RAG Quality & Retrieval Accuracy Testing

Test retrieval precision, source grounding, access control, hallucination risk, and response reliability.

Agent Evals & Workflow Testing

Evaluate task completion, tool selection, tool input quality, instruction adherence, and workflow reliability.

Prompt Regression Test Suite Development

Detect failures when prompts, models, context, tools, or retrieval sources change.

AI Red Teaming & Safety Testing

Test for jailbreaks, prompt injection, unsafe outputs, data leakage, and adversarial behaviour.

AI Governance

AegisIQ AI Governance Workbench

Build AI inventory, risk classification, policy controls, approval workflows, and evidence models.

AI System Inventory & Classification

Create a living register of AI systems, owners, models, data sources, integrations, and risk levels.

AI Risk & Impact Assessment

Assess privacy, fairness, safety, explainability, business impact, and operational risk.

AI Governance Framework Design

Define policies, review boards, evidence requirements, decision rights, and accountability structures.

RAG Integration Governance

Govern retrieval sources, access permissions, lineage, grounding quality, and leakage risk.

ISO 42001 Readiness

Prepare AI management-system controls, records, policies, and operating evidence.

EU AI Act / DPDP / GDPR Alignment

Map AI systems and data practices to relevant regulatory expectations.

AI Security

Prompt Injection & Jailbreak Defence

Detect and reduce prompt injection, jailbreak, tool-abuse and unsafe-action risk in agents and LLM applications.

RAG Data Leakage Controls

Reduce unauthorised exposure through permission-aware retrieval, redaction, filtering, and testing.

AI Access Management & Data Classification

Align data classification, access policies, and model usage controls.

AI Security Incident Response Planning

Define playbooks, escalation paths, and recovery processes for AI-related incidents.

Application, API & Cloud Security Review

Assess code, APIs, architecture, and cloud environments for exploitable risk.

Threat Modelling & Architecture Review

Identify design-level risks before systems are built or released.

AI Monitoring & Operations

Model / Agent Monitoring & Drift Detection

Monitor quality, behaviour, drift, safety, cost, and performance after deployment.

AI Cost, Latency & Reliability Tracking

Track operational cost, latency, token usage, and reliability against agreed thresholds.

Evaluation Feedback Loops

Route production findings back into evaluation suites, prompts, and retrieval improvements.

ISO 27001 / ISMS Readiness

Build security governance, risk treatment, Statement of Applicability, Annex A controls, and audit evidence.

Supplier & Third-Party Risk Management

Assess vendors, platforms, and technology dependencies for security and compliance risk.

vCISO & Security Governance

Provide senior security leadership for risk, control maturity, board reporting, and audit readiness.

SOC 2 Type II Readiness

Prepare control evidence and operating cadence for SOC 2 readiness.

Cross-Framework Compliance Mapping

Reduce duplicated compliance work by mapping controls across ISO, SOC 2, DPDP, GDPR, and sector expectations.

Use cases

Typical engagements.

  • Pre-release validation of GenAI systems
  • AI governance framework for a listed company
  • RAG data leakage risk reduction
  • Prompt injection defence
  • ISO 42001 readiness
  • AI system inventory and risk classification
  • Security assurance for AI-enabled applications

How Captivolt delivers

The engagement, and what you are left with.

ENGAGEMENT SHAPE

Typically a governance or AI-QE gap assessment, followed by framework implementation and an ongoing assurance cadence.

What makes this different

  • Combines AI-QE, governance, and security
  • Converts AI governance into operating evidence
  • Strong fit for regulated and listed enterprises
  • Proprietary VeriCore and AegisIQ frameworks

What you receive

  • AI system register and risk classification
  • Evaluation suites, scorecards, and regression baselines
  • Governance workflows and evidence model
  • Security findings and remediation plan
  • Monitoring and escalation runbooks

The evidence

What you are left holding.

Five artefacts, laid out as the records they are. Four of them are structures rather than filled-in examples, because the values belong to your system, and a page arguing that untraceable claims are worthless has no business printing invented ones.

Evaluation scorecard

Every dimension carries a pass condition, and most of them are must-pass: a release cannot proceed while one fails.

  • Response qualitygraded against agreed answers for the cases in scopeMUST PASS
  • Groundednessevery claim traceable to a retrieved passageMUST PASS
  • Retrieval precisionno regression against the accepted baselineMUST PASS
  • Hallucination rateunsupported claims below the ceiling agreed at DefineMUST PASS
  • Policy adherenceno response that breaches a named policyMUST PASS
  • Tool-use accuracythe right tool, with inputs that validateMUST PASS
  • Prompt-injection resistancethe adversarial suite passes in fullMUST PASS
  • Task completionthe workflow finishes on in-scope casesMUST PASS
  • Safetyrefusals correct in both directions: it declines what it must and answers what it mayMUST PASS
  • Cost / latencywithin the ceiling agreed before the buildTRACKED
  • Human escalation qualityescalates with enough context for a person to actTRACKED
  • Regressionnothing that passed before fails nowMUST PASS
  • Driftmeasured against the last accepted baselineTRACKED

The dimensions and pass conditions are ours; the scores are your system’s, measured on your cases. We do not publish example numbers, because a number nobody can trace is the thing this page argues against.

Release gate

The conditions checked before a version ships, and the three answers the gate is allowed to give.

  • Every must-pass dimension passes on the build being released
  • The adversarial suite has run against this version, not an earlier one
  • No open finding above the severity agreed at Define
  • Evaluation evidence attached to the release itself, not summarised in a message
  • A named owner for the system in production, and a path to reach them
And the three answers it may give
  • Release: the evidence is attached to the record, and the baseline becomes the one monitoring compares against
  • Hold, with the failing dimension, the reason, an owner, and the condition that would clear it
  • Override: possible, and recorded as an exception with an accountable approver and an expiry date

A gate that can only say yes is a status meeting. These are the conditions; the outcome is whatever the evidence produces on the day.

Red-team example

Six probes we run against agentic and RAG systems, each paired with the control that is supposed to stop it.

  • Indirect prompt injection through retrieved content

    What failure looks like

    A document in the corpus carries instructions and the agent follows them, having treated retrieved text as direction rather than as data.

    The control designed to stop it

    Retrieved content is never executed as instruction, and a tool call is authorised by policy rather than by the model’s intent.

  • Permission escalation through retrieval

    What failure looks like

    One user surfaces content only another is entitled to see, because the index was searched with the service’s access rather than the caller’s.

    The control designed to stop it

    Retrieval inherits the caller’s permissions and is checked per request, not per index build.

  • Tool use beyond remit

    What failure looks like

    A tool outside the task is called, or called with inputs nobody supplied, or called more times than the work requires.

    The control designed to stop it

    An allow-list registry, schema validation on every input, and an authority boundary the model has no way to widen.

  • Exfiltration through the response or an outbound call

    What failure looks like

    Credentials, secrets or personal data reach the output, a log, or a request to an external system.

    The control designed to stop it

    Output filtering, an egress policy on every tool, and redaction at the retrieval boundary.

  • Refusal bypass by reframing

    What failure looks like

    A restricted request succeeds once it is rephrased: as fiction, as translation, as code, as a hypothetical.

    The control designed to stop it

    Policy evaluated outside the prompt, where a reframing has nothing to argue with.

  • Confident answers from superseded context

    What failure looks like

    A withdrawn policy or an expired price is quoted as current, with a citation that makes it look verified.

    The control designed to stop it

    Freshness and supersession checks on the corpus, and no answer the retrieved evidence does not support.

Public attack classes, named in full and described at the level of the class and its defence. Findings from an engagement belong to the client and stay there.

Governance record

What is held for each AI system, so an auditor’s question has an answer that already exists.

  • System identity and version
  • Business owner, and the team that operates it
  • Purpose, and the decisions it influences
  • Risk classification, and the rule that produced it
  • Data sources, permissions and lawful basis
  • Model and version, with its change history
  • Tools it may call, and the authority for each
  • Evaluation evidence, by run rather than by claim
  • Approvals: who, when, and against which evidence
  • Exceptions, each with an owner and an expiry date
  • Incidents, and what changed afterwards
  • Next review date

The record structure, as implemented in AegisIQ. Populated from your systems: a register we filled in with examples would be a worked fiction.

Monitoring result

The signals watched after release, and the escalation path when one of them moves.

  • Groundedness on sampled production traffic
  • Retrieval quality against the release baseline
  • Refusal and over-refusal rates
  • Tool error and retry rates
  • Latency and cost per task against the agreed ceiling
  • Drift since the last accepted baseline
  • Failures reported by the people actually using it
When one of them moves
  • A threshold moves: alert, with the dimension named rather than a generic page
  • Triage by a named owner, against the same scorecard used before release
  • Re-evaluate on the current build and the current data
  • Accept the new baseline, fix, or pause, with the pause condition written before it was needed

Thresholds are set per system at Define, against the baseline that release produced. That is why they are not printed here as figures.

Evidence

Where we have done this.

  • THINK
  • ASSURE

Enterprise AI Framework for an NSE-listed Company

Real anonymised engagement
Client context
An NSE-listed company required a board-credible framework to take AI from initiative to governed operating capability.
Challenge
AI activity was growing faster than the governance, accountability, and evidence structures needed to oversee it.
What Captivolt delivered
Governance framework · use-case intake workflow · risk classification · accountability model · evidence requirements · oversight cadence.
What changed
A listed company moved AI from scattered initiative to a governed operating capability its board can oversee.
  • AI governance workflow
  • Risk classification model
  • Use-case intake design
  • Evidence model
  • BUILD
  • ASSURE

Enterprise AI QE Architecture

Proprietary framework
Context
GenAI systems routinely pass demos and fail in production, because they are not tested like enterprise software.
Challenge
LLM, RAG, and agentic systems need evaluation disciplines that traditional QE does not provide.
What Captivolt delivered
Evaluation architecture · dataset design patterns · regression suite structure · scorecard model · monitoring approach.
What it provides
AI quality became evidence rather than opinion: one repeatable architecture for testing LLM, RAG and agentic systems before and after release.
  • AI-QE lifecycle
  • Evaluation scorecard (concept)
  • Regression suite design

Relevant accelerators

What carries this work.

  • VeriCore AI Evaluation Studio →LLM, RAG & AGENT EVALUATION

    It is the measurement side of the ASSURE lifecycle: stages 02 Evaluate through 04 Gate, and 07 Monitor.

  • AegisIQ AI Governance Workbench →AI GOVERNANCE & CONTROL EVIDENCE

    It is the record side of the ASSURE lifecycle: stage 05 Approve, and the governance record every other stage writes to.

Discuss Your AI Initiative.

A structured first conversation about what you are trying to build, govern, or scale.