FILED: 2026-06-08 · TOPIC: Field Reports · 8 MIN
AI Red-Teaming Methods Compared — Anthropic, OpenAI, DeepMind
Three frontier labs publish red-teaming methodology. The methods diverge enough that a reader who reads them carefully can form a working view of what each lab thinks pre-deployment safety evaluation actually requires. A comparative read of the published methodology documents, with notes on what each method does well, what each underspecifies, and what an audit-grade red-team would look like.
Red-teaming, in the AI safety register, is a labelled-vague term. It covers everything from adversarial prompt engineering by a research intern to formal capability evaluation by a contracted external team. Three of the frontier labs — Anthropic, OpenAI, and Google DeepMind — have published methodology documents detailed enough that an outside reader can form a real comparative view. This piece is that read, focused on what each lab’s published method actually does and where the methods diverge in ways a procurement-side reader should care about.
We have read the published material. We have not characterized internal procedures the labs have not put on the record.
The shared core
The three published methods share enough of a core that the comparative work is meaningful.
All three split their pre-deployment evaluation into two regimes. The first is capability evaluation: structured testing of what the model can do, on tasks selected because they correspond to a class of risk the lab has identified. The second is adversarial elicitation: open-ended attempts to provoke the model into producing outputs the lab has committed not to ship. The capability evaluation is closer to a benchmark; the adversarial elicitation is closer to penetration testing.
All three publish a representative summary of the evaluations the model was subjected to before deployment. The summaries differ in detail and in candor.
All three carry a published commitment that the evaluation results condition the deployment decision. The strength of the commitment varies. The Anthropic Responsible Scaling Policy is the most explicit binding. The DeepMind Frontier Safety Framework is the most explicit about cadence. The OpenAI Preparedness Framework is the most explicit about decision process.
These shared properties define what “frontier-lab red-teaming” now means as a category. The interesting comparative work begins at the level of specifics.
Anthropic — the constitutional and the systematic
The Anthropic published methodology, principally documented in the Responsible Scaling Policy and in supplemental safety research papers, treats red-teaming as an extension of the lab’s constitutional-AI training methodology. The lab maintains an internal taxonomy of harm categories. For each category, the lab maintains structured prompts designed to elicit the corresponding harm. The evaluation phase runs the prompts against the candidate model. The results are scored against a published rubric. The scores feed into the deployment decision.
The methodology’s strength is its systematicity. The category taxonomy is published; the prompt families are described; the rubric is named. A reader of the published material can form a clear view of what the lab considers a covered harm category and what it does not. The published artefacts compose: the Responsible Scaling Policy specifies the deployment-binding commitment, the supplemental papers describe the methodology, and the model-card releases summarize the results.
The methodology’s limit is its distributional coverage. The structured prompts are good at the category of harm they are designed to elicit. They are less good at the category of harm the lab has not yet anticipated. The published methodology addresses this through periodic taxonomy updates and through complementary open-ended adversarial elicitation. The structured-and-systematic core is the published primary methodology; the open-ended supplements are how the lab handles distribution drift.
The methodology’s procurement implication is that an Anthropic-published model card’s safety results carry the lab’s structured-taxonomy framing. A buyer’s general counsel reading the card should understand the framing — the card reports performance on the categories the lab evaluates, not performance on every possible harm category the buyer’s deployment might encounter.
OpenAI — the cross-functional and the operational
The OpenAI published methodology, principally documented in the Preparedness Framework and in the lab’s system cards, treats red-teaming as a cross-functional operational process. The framework names the internal Preparedness team, the Safety Advisory Group, and the deployment-decision pathway through which the team’s findings travel. The methodology’s emphasis is on the process by which evaluation results convert into deployment decisions, more than on the evaluation methodology itself.
The methodology’s strength is its operational legibility. A reader can form a clear view of who, inside the lab, is responsible for what. The cross-functional structure is published. The decision pathway is published. The reader who wants to understand the institutional locus of the deployment decision can read the answer.
The methodology’s limit is its evaluation specificity. The published material is less detailed than the Anthropic material on the specific evals run, the specific prompts used, the specific scoring rubric. The released system cards summarize the evaluations at a higher level of abstraction. The reader who wants the methodology in detail has to infer more from the summary; the inference is plausible but is doing more work.
The methodology’s procurement implication is that an OpenAI-published system card carries the lab’s process framing. The card reports the outcome of the deployment decision the cross-functional structure produced. A buyer reading the card should understand that the card is the lab’s published outcome statement, not the lab’s published methodology.
DeepMind — the eval-cadence and the threshold-binding
The DeepMind published methodology, principally documented in the Frontier Safety Framework and in supplemental research papers, treats red-teaming as the data collection that feeds the Critical Capability Level evaluations. Each CCL has a defined eval suite. The eval suite is run on the model at defined cadence. The results determine the CCL tier the model has reached. The CCL tier determines the lab’s mitigation commitments.
The methodology’s strength is its threshold definition. The CCL-to-mitigation mapping is explicit. The evaluation cadence is explicit. The reader who wants to know what condition would trigger which mitigation can read it. The framework is, on this dimension, the most pre-committed of the three.
The methodology’s limit is the internal eval-suite specificity. The eval suites are described in summary; the full contents are not, in every case, public. The reader who wants to reproduce the eval against an open-weight comparison model has, for some CCLs, an incomplete picture. The reader who wants to verify a threshold has crossed has to trust the lab’s reported score.
The methodology’s procurement implication is that the DeepMind-published model card carries the lab’s threshold-based framing. The card reports the most recent CCL evaluation. A buyer reading the card should understand that the card is the lab’s published threshold position, with the corresponding mitigation commitments in force.
What each method does well, side by side
A working summary of the comparative position, on the published material:
- Methodology specificity. Anthropic is the most specific about what the evaluation actually does. The structured taxonomy and the published rubric are the most detailed of the three.
- Process specificity. OpenAI is the most specific about how the evaluation feeds into the deployment decision. The cross-functional structure is the most documented of the three.
- Threshold specificity. DeepMind is the most specific about what the evaluation outcome commits the lab to. The CCL-to-mitigation map is the most explicit of the three.
A procurement-side reader who needs all three things — the methodology, the process, and the commitment — has to read all three frameworks. No single lab’s published material covers all three at the level of detail the comparative reader gets from the union.
What each method underspecifies
The harder question is what each method does not cover that an audit-grade methodology would.
All three methods underspecify post-deployment monitoring. The frameworks address pre-deployment evaluation. The published material on what the lab evaluates after deployment — in response to which signals, on which cadence, with which intervention thresholds — is thinner. The pre-deployment regime is the published regime; the post-deployment regime is implied. An audit-grade method would publish both at comparable specificity.
All three methods underspecify agentic-system evaluation. The published methods are designed against single-model evaluation. The deployed systems the labs sell are increasingly agentic — multi-step, tool-using, orchestrated. The published methods do not yet say, in detail, how the single-model evaluations compose into the multi-step evaluation an agentic deployment would require. The bridge from per-model evals to per-system evals is one of the field’s open methodological problems.
All three methods underspecify third-party reproduction. The frameworks describe the methods the lab uses internally. They do not, in detail, describe how a third party — a regulator, an auditor, a peer lab — could reproduce the evaluation on the same model. The reproducibility question is not addressed in the published material. The published methods are, on a strict reading, the labs’ descriptions of what they did, not the labs’ specifications of how to check it.
All three methods underspecify counterfactual mitigations. The published material describes the mitigations the lab chose. It does not record the mitigations the lab considered and rejected, or the cost-benefit reasoning behind the choice. The published frameworks are commitment artefacts; they are not full publications of the lab’s safety-engineering reasoning.
What an audit-grade red-team would look like
A method that addressed the underspecifications above would, at minimum, do four things.
It would carry a published eval suite at reproducible specificity. The methodology document would include enough detail that a peer lab or an independent auditor could run the same eval on the same model and reach the same score. The result would be checkable.
It would carry a published post-deployment monitoring regime. The method would describe, at the same level of detail as the pre-deployment evaluation, what the lab monitors after deployment, what would trigger intervention, and what the published monitoring outcomes have been to date.
It would carry an agentic-evaluation extension. The method would describe how the lab evaluates the multi-step, tool-using agentic deployments built on top of the model, on the same kind of structured taxonomy the model-level method uses.
It would carry a published reasoning artefact. The method would publish, in the same register as the published commitments, the lab’s reasoning about why the chosen mitigations and thresholds are the right ones. The reader would not have to infer the reasoning from the conclusion.
These four are not, on the published frameworks, anyone’s current methodology. They are the direction the field will need to move if the published red-team artefact is to do procurement work commensurate with the deployments the artefact is supposed to cover.
The published frameworks are honest first attempts. The next-generation versions will be the test of whether the labs convert their published commitments into a methodology a third party can run and a regulator can audit. We will keep reading.
Methodology. This piece is sourced from the Responsible Scaling Policy, the Preparedness Framework, the Frontier Safety Framework, and from the labs’ published model cards and research papers. We have linked the underlying documents where we have characterized a position. We have not invented any methodology specifications.