AI Capture the Flag: From Hacking AI to the Human vs AI Cyber Exercise image

AI Capture the Flag: From Hacking AI to the Human vs AI Cyber Exercise

Oct 2026

|

8 min read

Capture the flag (CTF) is cybersecurity's hands-on game: competitors find and exploit weaknesses in deliberately vulnerable systems to retrieve hidden tokens called “flags”. As AI has become part of both cyber systems and cyber operations, it has also changed how CTFs are being used.

What is an AI capture the flag?

The term AI capture the flag can refer to two distinct types of exercise. In one, AI is the target: participants attack an AI system to expose weaknesses or make it behave in ways it should not. In the other, AI is the competitor: AI agents take on CTF challenges themselves, sometimes competing directly against human teams. Security teams are beginning to use both formats to test the security of deployed AI systems and to compare how humans and AI perform on the same cyber challenges.

AI as the target

In this format, the target is usually a large language model (LLM) or an application or agent built on one. Participants may try to bypass a guardrail, extract protected information or manipulate the AI into misusing a tool it can access. Successfully completing the objective reveals a flag that proves the challenge has been solved. This is the most common use of the term AI CTF and the format most AI CTF platforms are built around.

AI as the competitor

In the second format, AI solves the CTF challenges rather than being attacked. Entrants can range from fully autonomous agents that plan and carry out every step themselves to human AI teams in which a person guides the agent. They can compete against human teams on the same challenges, scoreboard and scoring rules.

CTFs that target AI systems

As companies deploy chatbots and AI agents in live products, CTFs have appeared that attack them. In many LLM-focused challenges, the attack happens through inputs rather than binary exploitation: participants craft prompts, or content the model will later read, to override instructions, extract protected information or steer an agent into misusing a tool it can access. Other challenges go beyond prompting and target the model's training data, the model itself or the surrounding application. These challenges fall into six broad types that draw on threat classes covered by the OWASP Top 10 for Large Language Model Applications and MITRE ATLAS, a catalogue of adversarial AI techniques.

· Prompt injection: overriding a model's instructions with crafted input, either directly or through content it later reads. This is the signature AI-CTF challenge.

· Jailbreaking and guardrail bypass: defeating the safety filters meant to block prohibited output.

· Secret and data extraction: coaxing out the hidden system prompt, a protected string, or data the model should never reveal.

· Tool and agent abuse: manipulating an agent's goals or getting it to misuse the tools and permissions available to it.

· Insecure output handling: getting the model to produce content that the surrounding application trusts and acts on, allowing the payload to reach a downstream system rather than remaining confined to the model.

· Classic machine-learning attacks: evasion, data poisoning, model inversion and model theft.

Head-to-head: when AI competes against humans

AI now competes in ordinary CTFs on the same scoreboard as human teams, under the same time limits and scoring rules. Everyone competes for the same leaderboard positions. Results from 2025 and 2026 show a consistent pattern: AI out-scored most of the human field, while human teams still held the top spot.

The first public head-to-head, the “AI vs Human” CTF run by Hack The Box with Palisade Research in 2025, put eight autonomous agents against 403 human teams over 48 hours. Five of the agents solved 19 of 20 challenges, yet the overall runner-up was a human team.

In 2026, NeuroGrid scaled the experiment up to 72 hours, with 1,337 human-only teams and 156 AI teams on the Hack The Box platform. Among teams that attempted a challenge, 73% of AI teams solved at least one, compared with 46% of human-only teams. Across all participants, the AI solve rate was 3.2 times higher. Among the top 5%, the gap narrowed to 1.69 times. The best human team still finished ahead at the elite tier. However, NeuroGrid's AI teams used a person in the loop rather than running fully autonomously, so the result reflects human-AI pairing as much as autonomous AI performance.

That same year, the startup Tenzai reported that its AI had beaten 99% of 125,000 human competitors across six CTFs without ever taking first place. DARPA's AI Cyber Challenge, decided at DEF CON 33 in 2025 with a four-million-dollar prize, tested AI systems with no human competitors at all.

The results also show where AI performs well and where it still struggles.

· In recent head-to-head competitions, AI performed best on bounded, well-documented work: textbook cryptography, easy-to-medium reverse engineering and standard exploitation of a kind it has seen many times.

· AI struggled most with black-box exploration, dynamic and runtime work, novel “guessy” challenges with no training precedent, and full-spectrum operations. Across these tasks, it also struggled to recognise when its answers were wrong.

01_dividing_line

Where AI has closed the gap in CTFs, and where human skill still decides the outcome.

What human-vs-AI cyber exercises reveal

When running a human versus AI CTF, the key question is not simply whether humans or AI perform better, but how much of the work the AI is allowed to do independently. A 2025 study of agentic CTF teams compared three levels of autonomy and found a clear performance pattern. Human-led teams, where a person drove and the model assisted, solved an average of 2.7 challenges. Fully autonomous agents solved 5.5. Hybrid teams, where the agent worked independently but handed off to a person at set points, solved 7.7, outperforming both. The advantage was greatest for challenges that needed debugging and repeated interaction with the environment, work that automated workflows run through more systematically than a human-led team.

05_autonomy_levels

In the 2025 agentic-CTF study, hybrid human-and-AI teams outperformed both human-led and fully autonomous teams.

A separate 2026 study of human-AI collaboration in a live competition described the most effective style as “pair hacking”: the agent handles high-throughput exploration while the human makes occasional strategic interventions. In one case, three brief human interventions helped an agent solve a reverse-engineering challenge it had been stuck on, moving it from second place to first. The same competition found that AI literacy partly compensated for limited domain knowledge: participants who were strong at prompting but new to CTFs scored 62%, compared with 27% for novices. That result suggests that security knowledge and the ability to use AI effectively are becoming complementary skills.

The study also showed how human-AI collaboration can break down, which matters just as much for exercise design. Most teams that planned to work with the AI step by step instead delegated an entire challenge to it with a “solve this” prompt and waited for a solution. Lower-skilled users often fell into what the study called “answer shopping”: repeatedly regenerating outputs in the hope that one would contain a solution. That approach almost never worked. The main bottleneck was usually not the model’s reasoning, but how participants framed the task and provided context. Unless the exercise helps them improve those skills, the results will reflect the AI’s capabilities more than the team’s ability to use it effectively.

Running a human-versus-AI cyber exercise

The same human-versus-AI comparison is moving beyond CTFs into full live-fire exercises. In 2026, NATO's Locked Shields, the world's largest live-fire cyber defence exercise, introduced AI-enhanced attack scenarios including automated reconnaissance, adaptive malware and AI-generated phishing. Participating teams also began fielding AI defensive agents of their own. AI-driven attack and defence is also appearing in the cyber ranges organisations use for day-to-day training: isolated, virtualised environments that stand in for real infrastructure. Five design choices largely determine how useful the results are afterwards.

· Set the autonomy levels explicitly. Decide whether teams will use AI in human-in-the-loop, fully autonomous or hybrid mode. Because hybrid teams performed best in the 2025 agentic-CTF study, treat “pair hacking” as a deliberate exercise mode rather than an afterthought.

· Design challenges that require sustained, operational work. In recent competitions, single-shot puzzles tended to favour AI agents. Tasks requiring stateful interaction, multi-step workflows and careful inspection are where human judgement made the biggest difference, and they are harder to automate fully.

· Use fresh, unpublished content. When solutions are already public, a model may have encountered the task or its write-up before, making it harder to distinguish reasoning from recall.

· Field AI on both sides. Run AI red and blue teams alongside human teams so the exercise shows how automation changes both attack and defence, rather than only puzzle-solving.

· Make the results traceable. Require conversation logs, agent trajectories and timing data so you can reconstruct how each result was reached: where AI helped, where it hallucinated, how many attempts it took and what the compute cost was. Feed that evidence back into a readiness programme.

03_who_won_to_how

How CTF scoring shifts once AI is a participant: from a leaderboard to a readiness benchmark.

In practice, this kind of exercise uses a cyber range to host realistic, isolated AI and infrastructure targets, with autonomous agents operating on both sides alongside human participants. The technical setup is only half the task. The harder part is recording how each result was reached so that the first exercise produces a baseline you can compare against when you repeat it next year.

What it means for security teams

If this pattern continues, with AI out-scoring most competitors, elite human teams keeping the top spots and human-AI teams performing best, it changes who you hire and what you train them to do.

· Automation is taking over more routine CTF work. That weakens both the value of CTF results as a talent signal and the traditional development path through which junior practitioners build toward senior expertise.

· Senior judgement remains especially valuable. Recent competition results suggest that experienced people add the most value in novel exploitation, full-spectrum operations and recognising when the AI is wrong.

· Supervision is its own skill. Pair hacking only works when people can scope tasks, provide the right context and verify the AI's output. Those abilities can be trained and are becoming as important as security knowledge itself.

From exercise to readiness

AI has changed capture the flag in two separate ways. It has become a formidable competitor, and recent studies have found the strongest results when humans and AI work together. It has also become a target in its own right, with a whole class of CTFs designed to break it. These are different exercises with different lessons: they share a name, but they test different things. What they share is practical value for security teams. Both show how AI is changing cyber work and help prepare people for adversaries using the same tools. For most teams, the practical next step is to run either a human-versus-AI exercise or a CTF against their own AI systems and see where AI succeeds, where it fails and where human judgement still matters in their own environment. Because those capabilities keep changing as models improve, the value comes from repeating the exercise rather than treating it as a one-off, turning each result into a baseline you can track over time.

For most organisations, the value is not simply knowing that AI can win a CTF. It is understanding how AI performs against their own people and defences, and keeping that picture current.

CybExer provides cyber-range environments and exercise design for organisations that want to test AI systems and the people who defend them. If an AI capture-the-flag or a human-versus-AI exercise forms part of your training and assurance plans, contact our cyber range experts to discuss what an appropriate exercise could look like.

Frequently asked questions

What is an AI capture the flag?

A capture-the-flag challenge involving AI: either a CTF in which the target is an AI system (prompt injection, jailbreaks, model attacks), or a CTF in which AI agents compete to solve the challenges, often against humans.

Can AI beat humans at capture the flag?

Against most of the field, yes. Recent head-to-head CTFs have seen AI entrants out-score the majority of human competitors. Against the best, not yet: elite human teams have still taken first place in every head-to-head covered here. Collaboration studies also suggest that human-AI teams can outperform either on their own.

How do you run a human-versus-AI cyber exercise?

Set the autonomy level (human-in-the-loop, autonomous or hybrid), use fresh challenges that require sustained operational work, field AI on both attack and defence, and require traceable logs so you can reconstruct how each result was reached. In practice, this type of exercise runs on a cyber range that can host AI agents and human teams together.

Is human-plus-AI really better than either alone?

Recent studies suggest so. Hybrid “pair hacking” teams, with agents exploring and humans steering, solved more challenges than both human-led and fully autonomous teams when participants were trained to scope tasks and verify the AI's work rather than simply delegate to it.

Related Resources

All news
Types of Cyber Security Exercises: A Comprehensive Guide
Read more
cybersecurity training solutions
How AI is Shaping the Future of Cybersecurity Training Solutions
Read more
What’s the Process Behind Red/Blue Team Exercise?
Read more
All blogs