AI Assurance

Test AI‑enabled systems before you rely on them

Evaluate how AI models, agents and AI-enabled systems behave in your enterprise environment without risking live systems.

Cyber range and digital-twin technology in use in more than 60 countries

European Space Agency NATO Support and Procurement Agency European Defence Agency Estonian Defence Forces Armed Forces of Ukraine Kuwait Institute of Banking Studies University of Tartu Spire Solutions

Test AI in Context

AI can perform well in a benchmark, vendor demonstration or isolated test and behave differently once connected to the systems around it.

As AI gains access to data, tools, APIs, identities and operational systems – and the ability to take actions – testing the model alone cannot show what happens in operation.

CybExer uses cyber-range and Digital Twin technology to represent the systems and dependencies relevant to the test. This allows AI-enabled systems to be tested repeatedly under controlled conditions without exposing production infrastructure.

Only the systems, dependencies and conditions relevant to the test need to be represented.

Digital Twin environment
Under test
AI Model / Agent
Access layer
Data & Tools APIs & Applications Identity & Permissions
Operating environment
Connected Systems Dependencies Security Controls

Shown within a representative Digital Twin environment.

What You Can Assess

CybExer AI Assurance enables organisations to test AI-enabled systems, evaluate performance in cyber operations, examine human-AI interaction and compare different options under equivalent conditions.

INPUTS REACH PromptDataTools MANIPULATED AI Agent model · policy · actions OT Critical Services Confidential Data Identity & permissions decide what is possible

AI Models, Agents & Applications

Evaluate third-party, open-weight, self-hosted and specialised AI models, agents and applications under defined conditions. Test how they behave when inputs are manipulated, information is incomplete, dependencies change or adversarial activity is introduced.

TELEMETRY COVERAGE & LIMITS AI detection True positive Triaged & escalated Missed / mis-classified What the tool catches and what it does not

AI-Enabled Security Capabilities

Test AI used in detection, analysis, cyber threat intelligence, vulnerability assessment and response. Evaluate performance, coverage and limitations under realistic cyber conditions.

HAND-OFFS · ESCALATION · SAFEGUARDS Analyst Decision & oversight AI system Permissions & tools tasking · context output · recommendation ESCALATION / HUMAN INTERVENTION Where a person must stay in the loop

Human-AI Operations

Test how people, processes and AI-enabled systems work together. Examine decision-making, permissions, escalation paths, hand-offs and where human intervention or additional safeguards may be required.

AGENT CONFIGURATIONS OBSERVED OUTCOMES Agent A Model M1 + Harness H1 + Tools T1 Agent B Model M2 + Harness H2 + Tools T2 Agent Team Agent C + Agent D Equivalent conditions same scenarios Baseline Reduced risk Unsafe actions Where behaviour diverges

Comparative Testing

Run different models, agents or AI-enabled tools under equivalent conditions to compare behaviour, security, performance and operational trade-offs.

Testing can address issues ranging from prompt injection, data exposure and unsafe outputs to unintended or excessive actions, privilege misuse, tool or API abuse and combinations of weaknesses across the wider system.

A combination to watch for

The Lethal Trifecta

Three capabilities are reasonable on their own and dangerous together: access to private data, exposure to untrusted content and the ability to communicate externally.

An attacker who can place instructions in content the agent reads can then have it retrieve private data and send that data out. An agent holding all three will leak sooner or later. This combination is widely known as the lethal trifecta.

AI Assurance testing shows which of the three an agent actually holds in your environment, and whether your controls genuinely remove one.

ONE AGENT, THREE CONDITIONS DATA EXFILTRATED Privatedata Untrustedcontent Externalcommunication Any two of these is workable. All three is not.

A weakness is found. What happens next?

From AI Weakness to Operational Consequence

A technical finding alone does not tell you how much risk it creates in your environment.

A weakness, failure or unexpected behaviour can have very different consequences depending on the systems, data, permissions and tools available to the AI.

CybExer AI Assurance can trace the wider impact: what the AI can access, what actions become possible, whether security controls prevent or limit them and what the resulting operational consequence could be.

01AI Behaviour / Weakness
02Access & Permissions
03Action / Attack Path
04Security Controls
05Operational Consequence

This helps organisations focus on the risks that matter in the environment where the AI will actually be used.

How AI Assurance Works

01

Define

Define the AI-enabled system or capability, the assurance question, relevant test conditions and the evidence required.

02

Represent

Create the relevant environment and integrate the model, agent, application or AI-enabled capability being assessed.

03

Test

Run expected, failure and adversarial scenarios under controlled, repeatable conditions without putting production systems at risk.

04

Assess & Generate Evidence

Assess behaviour, weaknesses, control effectiveness and consequences. Generate evidence to support deployment, procurement, security, risk and governance decisions.

Relevant tests can be repeated after remediation or when models, prompts, tools, permissions, integrations or data sources change.

Flexible deployment: Testing can be delivered in hosted or on-premises environments according to customer security, confidentiality and deployment requirements.

Built on Proven Cyber Range and Digital-Twin Capability

Founded in Estonia, CybExer’s cyber ranges and digital twin technology for large, multinational training and testing exercises is in use in more than 60 countries. AI Assurance builds on this established capability and experience across defence, government, cybersecurity and space.

60+ countries using CybExer technology
CybExer team deploying cyber range infrastructure at a multinational exercise Defence · Government · Space

Free preconfigured environment

Try Building a Digital Twin for AI Assurance

Start with a free, preconfigured cyber environment and adapt it to reflect part of your own system.

Explore how a digital twin can bring the systems, dependencies and controls relevant to your use case into a controlled environment for AI Assurance testing, and how it can be reused as your testing needs evolve.

The environment is built using CybExer vLM (Virtual Lab Manager), our AI-assisted technology for building and reusing cyber range environments and digital twins.

CybExer vLM used to build and reuse cyber range environments and digital twins

AI Assurance Q&A

What is AI Assurance?

AI Assurance helps organisations understand how an AI-enabled system behaves under defined conditions and what risks or consequences may arise when it interacts with the environment in which it is intended to operate.

CybExer focuses on testing AI in context, including relevant systems, data, identities, permissions, tools, dependencies and security controls.

How is AI Assurance different from AI red teaming or penetration testing?

AI red teaming can identify ways an AI system can be manipulated or made to fail. Penetration testing can identify vulnerabilities in technical systems.

AI Assurance goes further by examining the wider operational consequence: if a weakness exists or AI behaves unexpectedly, what can that enable, which controls intervene and what could happen across the surrounding systems or operations?

Isn’t this what an AI sandbox does?

AI sandboxes provide controlled environments for testing AI models, applications or code.

CybExer AI Assurance goes further by representing the systems, data, identities, permissions, dependencies and security controls that determine what an AI-enabled system can access and do within its intended operating environment.

Do you need to recreate the entire production environment?

No. Only the systems, dependencies and conditions relevant to the assurance question need to be represented.

What happens to sensitive models and data?

Testing is designed around the security, confidentiality and deployment requirements of the system being assessed. The delivery and data-handling approach is agreed according to customer requirements.

Is AI Assurance a one-time assessment?

Not necessarily. AI risk can change when models, prompts, guardrails, tools, permissions, integrations or data sources change, or when new vulnerabilities emerge.

Relevant tests can be repeated to understand whether those changes affect behaviour or risk and whether remediation has worked.

Does AI Assurance prove that an AI system is safe or compliant?

No. AI Assurance generates evidence about behaviour, controls, limitations and potential consequences under defined test conditions.

That evidence can support deployment, procurement, security, risk and governance decisions, but does not certify that an AI system is safe or compliant.

Test AI in the Context That Matters

Discuss the AI-enabled system, operating conditions and evidence you need before deployment, procurement or wider operational use.