▍ BLACK BOX NOTES ▍

DeepMind's Frontier Safety Framework — How 'Critical Capability Levels' Work

Google DeepMind's Frontier Safety Framework defines a tiered evaluation regime for frontier model capabilities and binds the lab's deployment posture to thresholds it has committed to in public. A working note on the framework's published structure, what the tiers actually mean, and what the regime exposes that competing labs' frameworks do not.

X LinkedIn Mastodon Print

The Frontier Safety Framework Google DeepMind first published in May 2024 — and updated in subsequent editions through 2025 — is, at the time of writing, the most explicit public commitment by a frontier lab to a tiered evaluation regime for its own models. Comparable artefacts exist at OpenAI (the Preparedness Framework) and at Anthropic (the Responsible Scaling Policy). The DeepMind framework is the one whose tier definitions and deployment-binding commitments are written in the most concrete form. This note reads the framework as published, names what the tiers actually require operationally, and records what the regime exposes that the comparable frameworks do not.

The shape of the framework

The Frontier Safety Framework is organised around the concept of Critical Capability Levels (CCLs). A CCL is a threshold capability — a specific level of model performance, on a specific class of task, that would, if reached, raise the model’s risk profile to a level the lab has committed to handling under additional mitigations. The framework lists CCLs for several capability domains. The published categories include autonomy (the model’s ability to act independently in extended task sequences), biosecurity (capability to produce content materially uplifting bio-threat actors), cyber (capability to materially improve offensive cyber-operations), and machine-learning R&D (capability to materially accelerate AI research itself).

For each CCL, the framework commits the lab to two things. The first is evaluation cadence: the model’s capability on the relevant task class is evaluated on a published schedule and at minimum before each significant deployment. The second is mitigation commitment: at and above the CCL, the lab commits to specific deployment-side mitigations the lab has named publicly. The mitigations are categorised. The lower CCL tiers correspond to enhanced monitoring and access controls. The higher CCL tiers correspond to deployment restrictions and, at the top, to commitments not to deploy until further mitigation work is complete.

The framework is, in the published form, the lab’s pre-commitment to its own deployment behavior. It is not a regulator’s instrument. It is a self-binding artefact — a public statement that the lab has named the thresholds at which its deployment posture changes, and that the deployment posture is now contingent on the evaluation outputs rather than on subsequent management decisions.

How the tiers actually work

The mechanics of the tier system, on the published material, are the part that distinguishes the DeepMind framework from the comparable artefacts at the other labs.

The CCL definitions are written at a granularity that allows external verification of the evaluation result. A given CCL in the autonomy domain, for example, is defined as a specific score on a specific eval suite, run against a specific model version, with specific scoring criteria. The evaluation methodology is published. The eval suite, where it is internal, is described in enough detail that an external researcher with comparable access could attempt to reproduce the result. The threshold is a number. The number is, on the published commitments, the threshold at which the lab’s deployment posture changes.

The mitigation commitments at each tier are written in a similar register. The lower tiers commit the lab to enhanced internal review and to specific access controls. The higher tiers commit the lab to capability-specific deployment restrictions — for example, restrictions on which downstream developers can access the model, on which tool configurations the model can be deployed with, and on which user populations the deployed system can be exposed to. The highest tier, where it has been written, commits the lab to a deployment hold until further mitigation work is complete.

The cadence of the evaluation is written as a contractual commitment. The lab has stated, publicly, that models will be evaluated at the CCL level on a schedule that does not lag deployment. The framework names the conditions under which a model would be deployed without an up-to-date evaluation — the published answer is roughly “none.” This is the strongest part of the published commitment. It binds the lab to a behavior independent of the lab’s commercial pressure to ship.

What the framework actually exposes

The interesting part of reading the framework as an auditability artefact is what it actually exposes about the lab’s operating posture.

The framework exposes the capability claims. A lab that publishes its evaluation methodology, its eval suite descriptions, and its threshold definitions has committed to a public language for talking about its model’s capabilities. Subsequent communications about the model — marketing copy, capability claims, deployment announcements — can be read against the framework. A discrepancy between the lab’s marketing register and its CCL-tier statements is now visible to an external reader. The framework’s documentation creates the contrast.

The framework exposes the deployment timeline. The CCL-tier mitigation commitments are conditional on evaluation outputs. The eval outputs are published, on the framework’s schedule, in enough detail that the lab’s deployment timeline is, in principle, predictable from the published material. An external reader can read the most recent CCL evaluation, read the corresponding mitigation commitments, and form an expectation about the deployment posture the lab is committed to. The lab’s actual subsequent deployment behavior can be checked against the expectation.

The framework exposes the decision points. The CCL thresholds are the named points at which the lab has committed that the deployment posture changes. The fact that a threshold exists at a specific score on a specific eval is, in itself, a kind of audit primitive. A regulator, a procurement-side auditor, or a peer publication reading the framework can ask the lab specific questions about specific CCL results, and the lab has committed to answer in the framework’s register.

What it does not expose

The honest read on the framework also requires naming what the framework does not expose.

The framework does not expose the internal deliberation. The decision to set a specific CCL threshold at a specific score is, in the published material, the lab’s decision. The reasoning behind the choice — why this score and not the adjacent one — is published in narrative form but not in a form that a regulator could independently audit. The framework is a commitment to a behavior; it is not a published research result about why the behavior is the right one.

The framework does not expose the eval-suite contents fully. The published eval methodology is descriptive enough to be plausible. It is not, in every case, complete enough that an external researcher could reproduce the eval from the published description alone. Where the eval is described in summary, the description carries the lab’s framing of what the eval measures.

The framework does not expose the competing mitigation choices. The mitigation menu at each CCL is presented as the menu the lab has chosen. The framework does not record the mitigations the lab considered and rejected, or the cost trade-offs that produced the chosen menu. This is, on a strict reading, the lab’s prerogative; the framework is a deployment commitment, not a publication of the lab’s full safety-engineering reasoning.

How it compares

The Preparedness Framework at OpenAI and the Responsible Scaling Policy at Anthropic carry comparable structures. All three label tiered capability thresholds. All three commit to deployment-side mitigations conditional on the threshold reached. The differences are at the level of specificity.

The Anthropic RSP is, on the published material, the most explicit about the relationship between capability and mitigation: each AI Safety Level (ASL) is described as a commitment that binds the lab’s deployment until specific mitigation standards are met. The DeepMind framework is, by comparison, the most explicit about the evaluation cadence and the eval methodology. The OpenAI Preparedness Framework, in the iterations published to date, has been the least concrete about the specific eval suite and the most concrete about the cross-functional decision process inside the lab. All three are honest first attempts at a public capability-evaluation commitment. Each exposes a different part of the lab’s operating posture; each obscures a different part.

A reader doing comparative auditability research can read all three together. The DeepMind framework is the one to read first for the eval methodology. The Anthropic RSP is the one to read first for the deployment-binding commitment. The OpenAI framework is the one to read first for the cross-team operationalization.

What this means for procurement

A procurement-side reader who has to make a deployment decision about a frontier-lab API in 2026 should treat the published safety framework as one of the procurement artefacts. The framework supplies the lab’s public capability claims, the lab’s published mitigation menu, and the lab’s stated evaluation cadence. A vendor that has not published a framework of this kind is operating without the equivalent commitment. A vendor that has published one and whose subsequent behavior has diverged from the published material is a vendor whose behavior an auditor should examine more closely.

The framework is not a regulator’s instrument. It is the lab’s public commitment to its own behavior. The procurement-side reader’s question is the simple one: has the lab in fact behaved consistently with its published framework? The publication’s view is that, on the published commitments to date, the DeepMind framework’s eval-cadence commitments are the most checkable. They are the commitments we will continue to track.

Methodology. This piece is sourced from the publicly published versions of the Frontier Safety Framework, the Responsible Scaling Policy, and the Preparedness Framework. Where we have characterized a position we have linked the underlying document. We have not invented any tier definitions or mitigation commitments.

Copied