AI Assurance

Test AI‑enabled systems before you rely on them

Evaluate how AI models, agents and AI-enabled systems behave in your enterprise environment without risking live systems.

Cyber range and digital-twin technology in use in more than 60 countries

European Space Agency NATO Support and Procurement Agency European Defence Agency Estonian Defence Forces Armed Forces of Ukraine Kuwait Institute of Banking Studies University of Tartu Spire Solutions

What the Testing Tells You

CybExer AI Assurance provides technical and operational testing of AI-enabled systems in realistic, controlled environments.

It produces evidence about their behaviour, security, performance and operational impact to inform deployment, procurement and governance decisions.

AI Assurance Under Real Conditions

AI can perform well in a benchmark, vendor demonstration or isolated test and behave differently once connected to the systems around it.

As AI gains access to data, tools, APIs, identities and operational systems – and the ability to take actions – testing the model alone cannot show what happens in operation.

CybExer uses cyber-range and Digital Twin technology to represent the systems and dependencies relevant to the test across IT, OT and IoT environments. This allows AI-enabled systems to be tested repeatedly under controlled conditions without exposing production infrastructure.

Only the systems, dependencies and conditions relevant to the test need to be represented.

Digital Twin environment
Under test
AI Model / Agent
Access layer
Data & Tools APIs & Applications Identity & Permissions
Operating environment
Connected Systems Dependencies Security Controls

Shown within a representative Digital Twin environment.

Built on Proven Cyber Range and Digital-Twin Capability

CybExer’s cyber ranges and digital twin technology for large, multinational training and testing exercises is in use in more than 60 countries. AI Assurance builds on this established capability and experience across defence, government, cybersecurity and space.

60+ countries using CybExer technology
CybExer team deploying cyber range infrastructure at a multinational exercise Defence · Government · Space

What You Can Assess

CybExer AI Assurance enables organisations to test AI-enabled systems, evaluate performance in cyber operations, examine human-AI interaction and compare different options under equivalent conditions.

INPUTS REACH PromptDataTools MANIPULATED AI Agent model · policy · actions OT Critical Services Confidential Data Identity & permissions decide what is possible

AI Models, Agents & Applications

Evaluate third-party, open-weight, self-hosted and specialised AI models, agents and applications under defined conditions. Test how they behave when inputs are manipulated, information is incomplete, dependencies change or adversarial activity is introduced.

TELEMETRY COVERAGE & LIMITS AI detection True positive Triaged & escalated Missed / mis-classified What the tool catches and what it does not

AI-Enabled Security Capabilities

Test AI used in detection, analysis, cyber threat intelligence, vulnerability assessment and response. Evaluate performance, coverage and limitations under realistic cyber conditions.

HAND-OFFS · ESCALATION · SAFEGUARDS Analyst Decision & oversight AI system Permissions & tools tasking · context output · recommendation ESCALATION / HUMAN INTERVENTION Where a person must stay in the loop

Human-AI Operations

Test how people, processes and AI-enabled systems work together. Examine decision-making, permissions, escalation paths, hand-offs and where human intervention or additional safeguards may be required.

AGENT CONFIGURATIONS OBSERVED OUTCOMES Agent A Model M1 + Harness H1 + Tools T1 Agent B Model M2 + Harness H2 + Tools T2 Agent Team Agent C + Agent D Equivalent conditions same scenarios Baseline Reduced risk Unsafe actions Where behaviour diverges

Comparative Testing

Run different models, agents or AI-enabled tools under equivalent conditions to compare behaviour, security, performance and operational trade-offs. The report places the results side by side so procurement teams can compare the options directly. It also records any failures and their possible effects on connected systems.

Testing can address issues ranging from prompt injection, data exposure and unsafe outputs to unintended or excessive actions, privilege misuse, tool or API abuse and combinations of weaknesses across the wider system.

AI Assurance by Sector

The risks change with the purpose and setting of the AI system. CybExer uses the same platform to focus each test on the failures that matter in a particular sector.

Atom-Services-card-finance
AI ACTION CONTROL OUTCOME AI Agent payments · accounts Fraud checksApproval limitsBlocked Unauthorised payment CONTROL BYPASSED

Financial services

In financial services, testing covers unauthorised payments, attempts to bypass fraud controls, exposed customer data or AI-assisted account takeover. Other scenarios include manipulation of trading or investment decisions, misuse of agent privileges, leakage between customers through RAG systems, and attacks on payments or APIs.

Financial services
N66A9310-2
RESTRICTED DOMAIN OPERATIONAL DOMAIN Classified recordsMission systemReporting toolsShared workspace AI agent CROSS-DOMAIN LEAKAGE

Government and national security

For government and national security systems, testing covers exposure of sensitive or classified information, privilege escalation, manipulation of mission systems, supply-chain attacks, leakage across security domains, misuse of autonomous agents or disruption to critical services.

Government and national security
f48061fd382b5485329e650078031a21
IT ZONE OT ZONE AI Agent tools · APIs Operator console SEGMENTATION Control interfacePLC / actuator Unsafe operation

Critical infrastructure

For critical infrastructure, the test looks at whether AI-initiated actions can reach OT systems, trigger unsafe actions, disrupt operations or compromise control-system interfaces.

Critical infrastructure
2c1a7dfc1ce045954365ad28f9bdeb75-min
APPLICATIONS & AGENTS INTERNAL ASSETS AssistantAgent APIBuild pipelineSource codeCustomer data Model & tools third-party · hosted IP leakage THIRD-PARTY AI RISK

Enterprise and technology

Enterprise and technology testing covers AI applications, agents and APIs, along with intellectual-property leakage, third-party AI risk and the use of AI in software development.

Three Levels of AI Assurance

Most AI testing looks for known weaknesses. That matters, but it is only the first of the three levels described here.

01

Deterministic Assurance — Is a known weakness present?

Tests cover prompt injection, jailbreaking, disclosure of sensitive data, manipulation of retrieval-augmented generation (RAG), misuse of agents and tools, weaknesses in identity and authorisation, and vulnerabilities in APIs or infrastructure.

02

Enterprise Adversarial Scenarios — What happens when weaknesses are combined?

At this level, several weaknesses are combined into realistic attack scenarios. These may involve multiple stages or AI agents, privilege escalation, manipulation of business processes, long-running autonomous behaviour or movement across connected systems. A weakness that looks minor on its own can become serious as part of a larger attack.

03

Frontier Vulnerability Research — What has not yet been found?

AI-assisted vulnerability research, hypothesis generation and controlled validation in isolated environments, used to surface attack paths that are not yet on anyone's list. Validated discoveries become permanent regression tests that run again as the system changes.

The lethal trifecta below is one example of an enterprise attack scenario.

A combination to watch for

The Lethal Trifecta

Three capabilities are reasonable on their own and dangerous together: access to private data, exposure to untrusted content and the ability to communicate externally.

An attacker who can place instructions in content the agent reads can then have it retrieve private data and send that data out. An agent holding all three will leak sooner or later. This combination is widely known as the lethal trifecta.

AI Assurance testing shows which of the three an agent actually holds in your environment, and whether your controls genuinely remove one.

ONE AGENT, THREE CONDITIONS DATA EXFILTRATED Privatedata Untrustedcontent Externalcommunication Any two of these is workable. All three is not.

A weakness is found. What happens next?

From AI Weakness to Operational Consequence

A technical finding alone does not tell you how much risk it creates in your environment.

A weakness, failure or unexpected behaviour can have very different consequences depending on the systems, data, permissions and tools available to the AI.

CybExer AI Assurance can trace the wider impact: what the AI can access, what actions become possible, whether security controls prevent or limit them and what the resulting operational consequence could be.

01AI Behaviour / Weakness
02Access & Permissions
03Action / Attack Path
04Security Controls
05Operational Consequence
06Assurance Decision

This helps organisations focus on the risks that matter in the environment where the AI will actually be used.

How AI Assurance Works

01

Define

Define the AI-enabled system or capability, the assurance question, relevant test conditions and the evidence required.

02

Represent

Create the relevant environment and integrate the model, agent, application or AI-enabled capability being assessed.

03

Test

Run expected, failure and adversarial scenarios under controlled, repeatable conditions without putting production systems at risk.

04

Assess & Generate Evidence

Assess behaviour, weaknesses, control effectiveness and consequences. Generate evidence to support deployment, procurement, security, risk and governance decisions.

Where Testing Takes Place

Testing can run entirely on your own infrastructure, even where there is no external network connection. Models, prompts, data, configuration and system details stay inside your environment.

CybExer can host the test environment instead. The technology is the same in either case and supports commercial, government, regulated and sovereign settings.

The service uses the same cyber-range technology that defence and government customers already use on premises and in isolated environments.

DSC05218
Your infrastructure NO EXTERNAL CONNECTION Digital twin environmentAI model or agent under testSecurity controls & dependencies SAME TECHNOLOGY CybExer-hosted Digital twin environmentAI model or agent under testSecurity controls & dependencies

Evidence for Governance and Compliance

Most governance and regulatory regimes expect an organisation to test an AI-enabled system against realistic misuse before relying on it. CybExer AI Assurance produces evidence based on what the system actually did under adversarial conditions, rather than on a questionnaire or the supplier's own claims.

EU AI Act

High-risk systems must be tested for accuracy, robustness and cybersecurity, including against foreseeable misuse, before they are put into service. CybExer provides the test conditions, results and records for inclusion in your technical documentation.

ISO/IEC 42001

An AI management system must show that risks have been identified and controls are effective. The test results support the impact assessment and show how the controls performed.

NIST AI RMF

The framework calls for AI risk to be measured, not simply described. Repeatable tests provide documented results that can be compared over time.

DORA and NIS2

When an ICT service covered by these rules uses AI, test results can support risk management, resilience testing and third-party assurance.

Internal governance

Model approval boards and procurement teams can use the same evidence for deployment gates, vendor due diligence and approval decisions, even where no specific regulatory framework applies.

The framework that applies to you is agreed before testing begins, so CybExer can shape the evidence around your organisation's obligations.

isa-scoreboard-wall
01 Baseline First pre-production test 02 Change New model, prompt or tool 03 Re-run Same scenarios, kept twin 04 Compare Differences against baseline Digital twin retained between tests

Ongoing Testing as the System Changes

AI risk does not stay fixed. A new model version, an edited system prompt, another tool, wider permissions or a different data source can change behaviour that had already been tested and accepted.

The digital twin is kept for future tests. When the system changes, the same scenarios can be run again and compared with the previous results. There is no need to rebuild the environment from scratch.

The first pre-production test establishes a baseline. Later tests make drift and other unexpected changes easier to spot by comparing new results with that baseline.

Testing can become part of the release process. Model updates, permission changes and new integrations are checked against earlier scenarios before they reach production. Any new vulnerability is added to the test set, so later engagements build on what has already been learned.

The first engagement takes the most work because it establishes the environment, test conditions and baseline. Once that foundation is in place, later tests can be more focused and faster. Assurance continues without starting over each time.

Free preconfigured environment

Try Building a Digital Twin for AI Assurance

Start with a free, preconfigured cyber environment and adapt it to reflect part of your own system.

Explore how a digital twin can bring the systems, dependencies and controls relevant to your use case into a controlled environment for AI Assurance testing, and how it can be reused as your testing needs evolve.

The environment is built using CybExer vLM (Virtual Lab Manager), our AI-assisted technology for building and reusing cyber range environments and digital twins.

CybExer vLM used to build and reuse cyber range environments and digital twins

AI Assurance Q&A

What is AI Assurance?

AI Assurance helps organisations understand how an AI-enabled system behaves under defined conditions and what risks or consequences may arise when it interacts with the environment in which it is intended to operate.

CybExer focuses on testing AI in context, including relevant systems, data, identities, permissions, tools, dependencies and security controls.

How is AI Assurance different from AI red teaming or penetration testing?

AI red teaming can identify ways an AI system can be manipulated or made to fail. Penetration testing can identify vulnerabilities in technical systems.

AI Assurance goes further by examining the wider operational consequence: if a weakness exists or AI behaves unexpectedly, what can that enable, which controls intervene and what could happen across the surrounding systems or operations?

Isn’t this what an AI sandbox does?

AI sandboxes provide controlled environments for testing AI models, applications or code.

CybExer AI Assurance goes further by representing the systems, data, identities, permissions, dependencies and security controls that determine what an AI-enabled system can access and do within its intended operating environment.

Do you need to recreate the entire production environment?

No. Only the systems, dependencies and conditions relevant to the assurance question need to be represented.

Does this cover OT and industrial systems?

Yes. A digital twin can include IT, OT and IoT systems. The test follows an AI-initiated action towards a control-system interface, recording the response of each security control and the likely outcome if those controls fail.

What happens to sensitive models and data?

Testing is designed around the security, confidentiality and deployment requirements of the system being assessed. The delivery and data-handling approach is agreed according to customer requirements. If needed, the entire test can run on your infrastructure with no connection to external networks. Models, prompts, data and system details never leave your environment.

Is AI Assurance a one-time assessment?

Not necessarily. AI risk can change when models, prompts, guardrails, tools, permissions, integrations or data sources change, or when new vulnerabilities emerge.

Relevant tests can be repeated to understand whether those changes affect behaviour or risk and whether remediation has worked.

Does AI Assurance prove that an AI system is safe or compliant?

No. AI Assurance generates evidence about behaviour, controls, limitations and potential consequences under defined test conditions.

That evidence can support deployment, procurement, security, risk and governance decisions, but does not certify that an AI system is safe or compliant.

Test AI in the Context That Matters

Discuss the AI-enabled system, operating conditions and evidence you need before deployment, procurement or wider operational use.