ENISA Backs Cyber Range Testing as the AI Act Gets Teeth image

ENISA Backs Cyber Range Testing as the AI Act Gets Teeth

Sep 2026

|

5 min read

Two dates, one message

In July 2026, the European Union Agency for Cybersecurity published ENISA's view on Cybersecurity in the Frontier AI Era, its assessment of what frontier AI models mean for defenders, manufacturers and national authorities. On 2 August 2026, the European AI Office gained enforcement powers under the EU AI Act over general-purpose AI models with systemic risk.

Read separately, these are a policy paper and a legal deadline. Read together, they mark the moment Europe moves from stating principles about AI security to demanding proof of it.

And proof needs somewhere to happen.

The problem ENISA describes

The paper's evidence is blunt. The window between vulnerability discovery and exploitation has collapsed from years to months and now, potentially, to minutes. Attackers can weaponise newly disclosed flaws within 15 minutes; the median time from initial access to data exfiltration has been compressed to 72 minutes. Researchers cited by ENISA describe the delta between discovery and weaponisation as "approaching zero."

The sharpest observation comes from CERT-EU: the danger of the latest models lies less in the volume of vulnerabilities they find than in their ability to chain findings across multiple steps, reason about application logic and produce exploitation paths that once required deep specialist knowledge, aka "what turns a list of individual flaws into a working attack." One industry example in the paper puts numbers on the shift: an organisation that logged around 80 CVEs in the first quarter of 2025 was close to 500 a quarter one year later, and around 500 per day once frontier-AI-enabled tools entered the picture.

Chained, environment-dependent attack behaviour carries a technical consequence. You cannot measure it with a static question set or a code scan. You have to watch the model operate somewhere.

Beyond training: a new job for a proven discipline

The somewhere already exists. Across defence, government, critical infrastructure and the private sector, cyber ranges have long been used to train cyber professionals, validate incident response and strengthen organisational resilience through realistic exercises using purpose-built environments where attack and defence can play out safely.

What frontier AI changes is who gets tested. Traditional benchmarks measure how a model performs against predefined tasks. Operational environments are messier: AI systems interact with users, enterprise infrastructure, security tooling and business processes while facing adaptive attackers and rapidly evolving threats. A range can recreate those interactions, letting an organisation observe how AI behaves as part of a wider operational system rather than in isolation and judge whether its outputs and actions can be relied upon under realistic cyber pressure.

That is not an academic concern. For defence and government organisations, financial institutions, critical infrastructure operators, space organisations and large enterprises, the risk is not only that an AI system may fail. It is that it may produce the wrong recommendation, miss a chained attack, or behave unpredictably at the moment decisions matter most.

Europe is now moving on both halves of that problem. The law setting the obligation, and its cybersecurity agency proposing what standardised evaluation could include.

What the AI Act now demands

Under Article 55 of the AI Act, providers placing general-purpose AI models with systemic risk on the EU market must assess and mitigate those risks. Those rules have applied to models placed on the market since 2 August 2025; providers of models already on the market before that date have until 2 August 2027 to bring them into compliance. What changed on 2 August 2026 is enforcement: the AI Office became empowered to investigate compliance, require mitigation measures, impose fines of up to 3% of global annual turnover or €15 million, whichever is higher, and, in extreme cases, restrict or recall models from the EU market.

The obligation is clear, and much of the compliance scaffolding already exists: the GPAI Code of Practice commits providers to identifying systemic risks (including large-scale cyberattack risks) applying safety and security measures across the model lifecycle, and reporting evaluation results to the AI Office, with Commission guidance filling in expectations. What remains unclear is more specific: how to test the cyber capabilities of a frontier model realistically, and how a regulator would compare one provider's evidence with another's. ENISA's paper says as much: adequate risk mitigations for models with advanced cyber capabilities, it notes, still need to be identified.

ENISA's proposal

That is the gap the paper's first European-level recommendation speaks to. Cautiously, as a direction rather than a decision:

"…it may be useful to establish EU-wide state-of-the-art benchmarks for the security evaluation of advanced AI models, including standardised testing against cyber ranges, exploitability metrics, and simulations of chained attacks, to set a consistent bar for assessing their capabilities."

Three components, one architecture: the environment (cyber ranges), the measurement (exploitability metrics) and the scenario (simulations of chained attacks, precisely the behaviour CERT-EU flagged as the real danger). Proposed, not prescribed. And one purpose: a consistent bar, because evaluation results only serve regulators if they are comparable across models, which is something that ad-hoc lab testing cannot deliver.

Three questions, one testing discipline

It helps to separate three questions that are related but distinct. Is the underlying model dangerously capable? That is the Article 55 question, aimed at frontier model providers, and the one ENISA's benchmark idea addresses most directly. Is a specific AI-enabled product or system secure? That is a different evaluation, and the one behind ENISA's call for national authorities to develop evaluation capacities for products with AI functionalities.

Another question would be: Is the organisation prepared when an AI system fails, errs or comes under attack? Here the paper urges defenders to run continuous simulations against live architectures, turning AI into an always-on capability for red teaming and defensive validation, and urges authorities to run AI-powered threat hunting and publish datasets from frontier AI simulations.

Different questions, different evidence. What they share is the underlying discipline: environments that are controlled, instrumented and repeatable, where behaviour can be observed rather than assumed. That is the definition of a cyber range.

What CybExer can test today

Realistic evaluation environments are what CybExer builds. On our AI cyber ranges organisations can run AI models through agentic attack-and-defence scenarios against replicated infrastructure, test AI-enabled security products against live attack traffic, and exercise their own teams and processes in a twin of their operational environment. Before those systems are relied upon in critical environments. Each run produces a record: actions taken, detections triggered or missed, time-to-detect and time-to-respond, a full audit trail — the raw material that AI assurance is made of.

To be precise about scope: that answers the system-level and operational-level questions today, and it gives model-level capability testing a realistic place to run as EU-wide benchmarks take shape. What those benchmarks will formally require is still being defined. Which is exactly why we are investing in this direction now, together with our customers, rather than waiting for the standard to arrive fully formed.

"Cyber ranges have long helped organisations understand how people and systems perform during cyber incidents. As AI becomes part of critical security operations, the same principle applies: trust should be built through realistic evaluation, not assumed. It is encouraging to see this reflected in ENISA's recommendations."

— Lauri Almann, Co-founder and Member of the Board, CybExer Technologies

The clock did not stop in August

Six weeks after AI Act enforcement began, on 11 September, ENISA took on an operational role under the Cyber Resilience Act as manufacturers begin mandatory reporting of actively exploited vulnerabilities through ENISA's new Single Reporting Platform, a platform the frontier AI paper itself says "must be leveraged, once it is functional." The course is increasingly clear: security in Europe is becoming something organisations demonstrate continuously, not something they claim once.

The EU-wide benchmarks ENISA proposes will take time to formalise. The agency says it will refine its recommendations with Member States and align them with the European Commission's Action Plan. Enforcement will not wait for that process. The organisations best placed for what comes next are those building evaluation muscle now, because they will help shape the standard rather than scramble to meet it.

The question, in other words, is no longer only where to use AI. It is how much access can safely be given to critical systems, operational processes and sensitive information, and whether it can be proven trustworthy while the environment around it is under attack. If you are working out what that proof should look like — for your models, your products or your security operations — that is exactly the conversation our range was built for. Talk to our team about testing these capabilities in your operational context.

 

Related Resources

All news
AI Assurance Platforms: What They Cover, and What They Can and Can't Test
Read more
What’s the Process Behind Red/Blue Team Exercise?
Read more
AI Fabric: Making AI a True Teammate in Cyber Ranges
AI Fabric: Making AI a True Teammate in Cyber Ranges
Read more
All blogs