A red/blue team exercise is a simulation of a cyberattack, where two teams, the red team and the blue team, are pitted against each other. The red team represents the attackers, while the blue team represents the defenders. The exercise takes place on a cyber range that simulates the real-world IT infrastructure that the organisation is looking to protect.
The exercise is a valuable training tool for IT professionals and security personnel to improve their skills and prepare for real-world cyber threats. By simulating a cyberattack in a controlled environment, participants can learn how to detect, respond, and mitigate cyber threats, improving the overall security of their organisation's IT infrastructure.
But a good exercise is more than an attack and a defence. It is a repeatable process, and most of its value is decided in the preparation, long before anyone touches a keyboard. End to end, that process runs in ten stages:
- Define objectives
- Set scope and rules of engagement
- Build the environment
- Design the scenario
- Assign teams and roles
- Brief the participants
- Execute the exercise
- Monitor and score performance
- Debrief
- Track remediation, then repeat
What Is the Goal of a Red/Blue Team Exercise?
The goal is to produce a factual account of how a defence performs against a thinking opponent. Everything else about the format follows from that. The red team exists to generate the attack, the scoring exists to measure the response, and the debrief exists to convert the result into changes.
Stated that way, the goal governs every decision downstream: which systems are in scope, how capable an adversary the red team emulates, what the blue team is told in advance, which scoring categories carry weight, and what counts as an acceptable result. When an exercise disappoints, the usual cause is that this question was never answered precisely, and the process was assembled around a date in the calendar instead.
How a Red/Blue Team Exercise Helps an Organisation
Because the process is documented from objectives through to remediation, the exercise leaves behind artefacts rather than impressions. Four of them do most of the work.
- A timeline of what actually happened. The red team’s actions are logged and timed, so the account of the incident is not reconstructed from memory afterwards. It is the input to everything below.
- A measured response rather than an estimated one. Time to detect, time to contain and service availability are recorded during play, which makes them comparable between runs and between teams.
- A remediation list with owners. Findings come out of the after-action review attached to specific controls and procedures, which is what makes them possible to close.
- Something to show people outside security. A board or a regulator asking whether the organisation could handle an intrusion is asking for more than a maturity rating. A scored exercise with an attack timeline attached answers the question in the terms it was asked.
There is a further effect that appears in none of the artefacts. A team that has worked an incident together, even a simulated one, moves faster the next time, because the sequence is familiar rather than novel.
Stages 1–2: Defining Objectives, Scope, and Rules
A red/blue team exercise is typically facilitated by an instructor who is responsible for planning and preparing the exercise, selecting the participants for each team, and providing guidance to the participants throughout the exercise. What "planning and preparing" actually involves is among the most consequential parts of the whole process, because these decisions determine whether the exercise teaches anything useful. Before scheduling anything, settle:
- Business and learning objectives What the exercise must prove or improve (e.g. faster detection, cleaner escalation, a validated playbook).
- Scope Which systems are in play, and which are explicitly off-limits.
- Target systems and protected assets The "crown jewels" the blue team must keep available.
- Threat-actor profile The adversary the red team will emulate, and how capable they are.
- Duration and schedule Length, phases, and any pauses.
- Participant skill levels So the difficulty fits the audience.
- Rules of engagement Permitted and prohibited actions for both teams.
- Safety and stop conditions What pauses or halts the exercise, and who can call it.
- Success criteria and scoring methodology How performance will be judged (see Stage 8 below).
- Communications plan How teams, control, and observers talk during play.
- Technical rehearsal and a known baseline A dry run before launch, plus a clean baseline state and a reliable reset procedure.
Stages 3–4: Building the Environment and Scenario
With objectives set, the environment is built. That environment is a cyber range, meaning virtualised copies of servers, networks and applications running in isolation from anything real, which is what lets the red team break things freely. A standardised range uses generic systems. A digital twin reproduces your own estate, its configurations and its tooling, so the findings apply to what you actually run. Either way it is built to mirror the infrastructure under study, captured as a baseline snapshot so it can be reset cleanly between runs. The scenario is then designed around the objectives: a storyline, the injects that drive it forward, and the specific conditions the blue team must handle. A short technical rehearsal confirms that telemetry, tooling, and scoring all work before real participants arrive.
Blue Team Scenarios and Simulations
The scenario is the specific situation the blue team is asked to handle, and it is chosen to exercise the parts of the response that matter to the objective. Because the scoring categories reward availability, reporting and recovery alongside defence, a well-chosen scenario forces the team to trade those against one another rather than optimise for a single one.
Scenarios that recur in blue team simulations include:
- Perimeter service exploitation. An internet-facing application or edge device is compromised. The team has to find the entry point and decide how much of the service to take offline while it does so, which sets detection directly against the availability score.
- Domain compromise. The red team reaches credentials with wide privileges. The question is whether the team notices the privilege escalation at all, and whether it can cut off access without locking the business out of its own systems.
- A destructive attack requiring restoration. Data or systems are damaged rather than stolen. This is the scenario that tests backups honestly, because restoring under time pressure is a different thing from believing a backup exists.
- Business email compromise leading to fraud. There is no malware and little to detect technically, so the exercise moves onto reporting and escalation, and onto whether anyone outside the security team hears about it in time.
- Third-party access abuse. The attacker arrives through a supplier connection that is legitimately trusted, which tests whether monitoring reaches the parts of the estate the organisation does not own.
Running these as simulations on a cyber range rather than as discussions is what removes opinion from the result. Either the team found it or it did not, either the service stayed up or it did not, and either the restore worked or it failed.
Stage 5: Assigning the Teams and Roles
The participants in the exercise are typically IT professionals or security personnel who are responsible for securing the organisation's IT infrastructure. Beyond the red and blue teams, a mature exercise usually separates responsibilities that a single "instructor" cannot realistically carry alone:
- Exercise director / white (control) team Owns objectives, adjudication, and the go/stop decision.
- Green team Builds and maintains the environment during play.
- Scoring team Applies the scoring model consistently.
- Observers and evaluators Capture what happened for the debrief.
- Purple-team facilitators Bridge red and blue when the goal is collaboration rather than competition.
- Business or executive participants Practise decisions and communications under pressure.
- Legal, safety, and communications representatives As the scenario requires.
In a small training exercise, one instructor may wear several of these hats. In a large or high-stakes exercise, they should be distinct roles, otherwise the same person is planning, controlling, coaching, scoring, and evaluating at once, which undermines objectivity. The stages below keep saying “the instructor”, which is the single-facilitator model most training exercises actually use. In a larger exercise, read it as the control team, with each specific job sitting with whichever role above owns it.
Stage 6: Briefing the Red and Blue Teams
At the start of the exercise, the instructor provides a briefing to both teams. The red team is briefed on the target organisation's IT infrastructure, vulnerabilities, and attack vectors, while the blue team is briefed on the red team's tactics, techniques, and procedures (TTPs).
To be clear about who knows what: neither team is told the other’s plan. The red team is given the environment and its weaknesses but not the defenders’ playbook, and the blue team is told the kind of adversary to expect, not the specific actions it will take or when. What both teams share is the business context and the rules of engagement.
Should the Blue Team Know About the Exercise in Advance?
Two questions get confused here. One is whether the blue team knows an exercise is running at all. The other is whether it knows what the red team intends to do.
For a red/blue team exercise the first is almost always yes. The teams are assembled, briefed and scored, so the exercise is announced by construction. That is what separates it from an unannounced red team engagement, where the defenders are told nothing and the object is to see how an ordinary day goes.
The second question is the real design decision. Telling the blue team the specific TTPs turns the exercise into a tuning session, in which the team knows roughly what to look for and practises finding it. That is the right choice when the objective is training or validating a detection rule. Withholding them measures something closer to genuine capability, at the cost of a slower exercise in which teams may spend hours on the wrong thing.
As a working default, brief both teams on the same business context and the same rules, give the red team its objectives rather than a map of known vulnerabilities, and tell the blue team that an exercise is running without telling it what to expect. Depart from that when the objective is explicitly training, where handing over the TTPs is the whole point of the session.
One constraint applies whichever way you decide. Somebody on the defensive side has to know the exercise is running and have the authority to stop it, otherwise a real incident arriving in parallel will be waved off as part of the game.
Stage 7: Executing the Attack and Defence
The red team then begins executing their attack, using various attack methods such as phishing, social engineering, malware, and exploitation of vulnerabilities.The blue team works to detect, respond, and mitigate the attack using tools and techniques such as intrusion detection systems, firewalls, and security information and event management (SIEM) systems.
In practice the two sides move in step. A representative sequence looks like this:
| Red-team activity | Blue-team activity |
|---|---|
| Reconnaissance | Establish a baseline and monitor telemetry |
| Initial access | Detect the alert and begin triage |
| Credential access | Investigate affected accounts |
| Lateral movement | Scope the incident and contain systems |
| Objective execution | Protect critical services |
| Persistence or exfiltration | Eradicate access and begin recovery |
| Evasion or evidence removal | Preserve evidence and report status |
Throughout the exercise, the instructor provides guidance and feedback to both teams, helping them to improve their performance and adapt to new challenges. The instructor may also introduce new scenarios or challenges to the exercise to test the participants’ ability to adapt and respond to evolving threats.How much intervention is appropriate depends on the mode:
- Coached
Participants receive hints and guidance in real time, which is best for building skills. - Assessment
Intervention is kept to a minimum so genuine capability can be measured. - Adaptive
The control team raises or lowers difficulty based on how the teams are performing.
Stage 8: Monitoring and Scoring Performance
Scoring is how an exercise turns activity into evidence. CybExer's platform, for example, breaks each blue team's performance into weighted categories on a live scoreboard:

Team scores breakdown & total, all teams. Each column represents each Blue Team's score breakdown in a given period for each of the currently applied scoring categories: Availability (green), Incident Reports (light blue), Situation Reports (dark blue), Attacks (red), Restores from Backup (yellow) and Special (orange). In addition, Overall Score (purple) is provided.
Those categories exist because they match what a defender is actually judged on during an incident: keeping services running, reporting accurately, and recovering what was damaged. Whichever model you apply, the weighting should follow the objectives set in Stage 1, and it should be clear on:
- What earns and loses points Keeping services available and filing accurate incident and situation reports earn points. Letting the red team succeed, or causing unnecessary downtime, costs them.
- Automated vs assessed Availability and successful attacks can be measured automatically. Report quality is judged by the scoring team.
- Weighting by business impact A hit on a crown-jewel service should count for more than a peripheral one.
- Partial success How a contained-but-not-prevented attack is scored.
- Whether the red team is scored too, and against what goals.
- Penalties For false positives, unsafe actions, or downtime the business would never accept.
- Anti-gaming So teams optimise for real security outcomes, not for the scoreboard.
Stages 9–10: Debriefing, Remediation, and Repetition
After the exercise is complete, the instructor conducts a debriefing session with the participants. During the debriefing, they discuss their observations, lessons learned, and areas for improvement. This feedback helps participants to learn from their mistakes and improve their performance in future exercises.
That discussion is where the value starts rather than where it ends. A rigorous after-action review also:
- Reconstruct the attack-and-response timeline side by side.
- Compare what the red team actually did against what the blue team could see, and surface the missed detections.
- Review the containment and recovery decisions, and the communication and escalation that surrounded them.
- Separate platform or scenario issues from genuine participant performance.
- Map findings to specific controls and procedures, assign remediation owners with deadlines, and then re-run the exercise to confirm the fixes worked.
It is that remediation loop, not the discussion, that turns an exercise into organisational improvement.
How Often Should an Organisation Run a Red/Blue Team Exercise?
The process is a loop rather than a line, so the interval between exercises matters as much as any single run.
A full red/blue team exercise, with a built environment and a scripted campaign, is commonly run once or twice a year. That spacing gives the remediation from one round time to be implemented and then tested in the next, which is the mechanism that earns back the cost. Running one every few years produces a snapshot and little else, because the findings go stale before anybody acts on them.
Some changes justify an exercise regardless of the schedule. A migration, a merger or a new business-critical service alters the estate the last exercise was run against. New detection tooling is unproven until somebody has watched it fire against a live attack. A change of personnel counts for more than it appears to, because response capability held in two people’s heads leaves when they do. Regulated sectors will have minimum testing frequencies set for them, and those are a floor rather than a target.
Conclusion
What the exercise leaves behind is a measurable picture of the organisation under attack: which defences worked, which did not, how long the response took, and what has to change before the next round. Run as a disciplined, ten-stage process, with real preparation, honest scoring and a remediation loop, that picture is the difference between believing you are prepared and being able to show it. To design an exercise around your objectives, talk to CybExer's cyber range experts.