Regulating AI · Playbook
Thank you for reading this post, don't forget to subscribe!
Home / Indexes / Regulating AI / Controlling AI Research
Controlling AI Research: A Best-Practices Playbook
If a law can’t reliably switch a model off, the real controls have to live upstream — in how frontier AI is researched, tested, and released. Seven mechanisms now do most of that work. This is the sourced playbook: what each one is, who actually does it, and where it’s strong or leaky — plus an honest look at the tensions no framework has solved.
Last updated: July 2026. Every mechanism links to a primary framework or authoritative source.
The playbook at a glance
These seven mechanisms are what “controlling AI research” actually means in practice today. Most are voluntary and lab-authored; the rightmost column is our read on how binding each one really is.
| # | Mechanism | What it controls | How binding today |
|---|---|---|---|
| 1 | Responsible scaling / frontier safety frameworks | Ties safeguards to capability thresholds | Voluntary, self-authored |
| 2 | Dangerous-capability evals & red-teaming | Detects CBRN, cyber, autonomy risk pre-launch | Voluntary, contractual access |
| 3 | Compute governance | Reporting thresholds + chip export controls | Partly binding (EU); US rolled back |
| 4 | Model-weight security | Stops theft/exfiltration of the model | Voluntary; RAND SL framework |
| 5 | Incident reporting & transparency | Discloses capabilities, safeguards, failures | Binding for EU systemic-risk models |
| 6 | Independent oversight & whistleblowers | Checks the labs from outside | Weakest leg; CA SB 53 a start |
| 7 | International coordination | Shared testing, evidence, and norms | Diplomatic, non-binding |
01 Tie safeguards to capability thresholds
The foundational practice: define capability levels, and commit to specific safeguards — and to not deploying — until each level’s protections are in place. Every major lab now publishes one. Anthropic’s Responsible Scaling Policy uses AI Safety Levels (ASL) modeled on biosafety; in May 2025 it activated ASL-3 protections alongside Claude Opus 4 — notably as a precaution, before it had confirmed the model crossed the CBRN threshold. OpenAI’s Preparedness Framework v2 (April 2025) gates on “High” and “Critical” capability in biology/chemistry, cyber, and AI self-improvement. Google DeepMind’s Frontier Safety Framework uses “Critical Capability Levels.” The shared template comes from the Seoul Frontier AI Safety Commitments, which 16 companies signed in May 2024.
Best practice: publish the thresholds in advance, specify who decides one has been crossed, and default to caution when uncertain.
02 Test for dangerous capabilities before release
Frameworks are only as good as the evaluations that trigger them. Best practice is structured, adversarial testing for uplift in CBRN/bio, offensive cyber, and autonomy/self-replication — done internally and by independent third parties, pre-deployment. METR measures how long a horizon of tasks a model can complete autonomously (a loss-of-control proxy); Apollo Research probes for deceptive “scheming”; the UK AI Security Institute and US counterpart run pre-deployment testing of frontier models.
03 Govern compute
Compute is the one input that’s measurable and chokeable, which makes it the favorite lever of hard regulation. The EU AI Act presumes “systemic risk” above 1025 FLOP and requires notifying the AI Office within two weeks. The former US Executive Order 14110 set a 1026 FLOP federal reporting line — but it was rescinded in January 2025, so that mandate no longer binds. On the hardware side, chip export controls and “know-your-customer” certification for large compute buyers are the enforcement teeth — though 2025 saw significant US rollback, showing how quickly the geopolitics can override the safety rationale.
04 Secure the model weights
Every other control assumes the developer still holds the weights. If they’re stolen, the full capability leaks with none of the guardrails. RAND’s five security levels (SL1–SL5) — from defending against amateurs to withstanding top-tier nation-state operations — are the reference framework, and labs now map their hardened “security standards” toward it. Best practice pairs weight security with a deliberate release decision: structured (API-only) access keeps monitoring, usage policies, and revocation on the table; open-weight release is legitimate and pro-competition but forecloses all of them, because safety fine-tuning can be cheaply stripped once weights are public.
05 Report incidents and publish safety cases
Transparency is where voluntary practice is becoming hard law. Labs publish system cards (per-model capability and safeguard documentation) and increasingly safety cases (structured arguments that a model won’t cause a specific catastrophe). The EU AI Act GPAI Code of Practice (final July 2025) makes serious-incident reporting mandatory for systemic-risk models — with deadlines as tight as two days for incidents hitting critical infrastructure — and full fines from August 2026. California’s SB 53 adds a US state-level critical-incident channel.
06 Build real outside oversight
This is the weakest leg — and the most important, precisely because it’s the one labs don’t control. Today’s governance is largely voluntary and self-authored: the labs write their own thresholds and decide when they’re crossed. Strengthening independent oversight means three things: whistleblower protection (California’s SB 53 now shields employees raising catastrophic-risk concerns), a safe harbor for good-faith safety researchers (proposed by Longpre et al. and 350+ signatories, after labs suspended researchers’ accounts), and accredited third-party audits — which barely exist yet compared to financial auditing.
07 Coordinate internationally
No single country can govern a global technology. The scaffolding: the summit series (Bletchley 2023 → Seoul 2024 → Paris 2025 → India 2026), the International Network of AI Safety Institutes (launched November 2024 with 10 founding members), and the Bengio-chaired International AI Safety Report (first published January 2025, ~96 experts) that gives everyone a shared evidence base. Best practice is joint pre-deployment testing and shared incident data — but coordination is fraying: the US and UK declined to sign the Paris statement, and China sits outside the institute network entirely.
The tensions no playbook resolves
- Innovation vs. precaution. The summits literally renamed themselves from “Safety” to “Action”; California’s sweeping SB 1047 was vetoed and replaced by the narrower SB 53. These controls are real but politically reversible.
- Open weights vs. control. Mechanisms 1–6 all assume the developer keeps the weights. Open release is genuinely good for competition and scrutiny — and forecloses those levers. There is no clean answer.
- US–China competition. Compute governance is simultaneously a safety tool and an industrial weapon; when they conflict, safety tends to lose.
- Who writes the rules. Most frontier governance is voluntary and lab-authored. That’s why the independent-oversight leg (practice 6) is the least developed — it’s the one the labs don’t get to write.
Frequently asked questions
How is frontier AI research actually controlled today?
Mostly through voluntary, lab-authored frameworks: responsible scaling policies, dangerous-capability evaluations, weight security, and transparency reporting — supplemented by binding EU compute and incident rules and a thin layer of international coordination.
What’s the single biggest gap in AI governance?
Independent oversight. Today the labs largely set their own safety thresholds and judge when they’re crossed. Accredited third-party audits barely exist, and researcher safe harbors remain proposals, not law.
Do safety rules stop open-source AI?
Largely no. Once model weights are released publicly, most control mechanisms — revocation, monitoring, shutdown — no longer apply, because the model can be copied and its safety fine-tuning removed. That is the central unresolved tension in AI governance.
Sources
• Anthropic, Activating ASL-3 Protections (May 2025) — anthropic.com
• OpenAI, Preparedness Framework v2 — openai.com
• Google DeepMind, Frontier Safety Framework — deepmind.google
• AI Seoul Summit, Frontier AI Safety Commitments — gov.uk
• METR (autonomy evaluations) — metr.org
• EU AI Act, GPAI systemic-risk threshold (Art. 51) — artificialintelligenceact.eu
• Federal Register, rescission of EO 14110 (Jan 2025) — federalregister.gov
• RAND, Securing AI Model Weights (SL1–SL5) — rand.org
• EU AI Act GPAI Code of Practice (incident reporting) — artificialintelligenceact.eu
• Longpre et al., “A Safe Harbor for AI Evaluation and Red Teaming” — arxiv.org/abs/2403.04893
• NIST, International Network of AI Safety Institutes — nist.gov/aisi
• International AI Safety Report (Bengio, 2025) — internationalaisafetyreport.org