Regulating AI · Rating System
Thank you for reading this post, don't forget to subscribe!
Home / Indexes / Regulating AI / Model Security Ratings
AI Model Security Ratings 2026
“Best model” leaderboards rank capability. This one ranks something else: how much security and control a frontier lab has built around its model — the framework, the testing, the weight protection, and the transparency. It is a separate scorecard, graded on what each developer publicly discloses, not on how smart the model is.
Ratings dated July 2026. The AI Index’s disclosure-based assessment — an editorial scorecard, not a certification.
The scorecard
Six security dimensions, letter-graded, plus an overall. Grades reflect public frameworks, system cards, third-party testing participation, and release model.
| Developer / model | Safety framework | Capability testing | Weight security | Transparency | Oversight | Overall |
|---|---|---|---|---|---|---|
| AnthropicClaude | A | A | A- | A | A- | A |
| OpenAIGPT | A- | A- | B+ | A- | B | A- |
| Google DeepMindGemini | A- | B+ | A- | B+ | B | B+ |
| MicrosoftCopilot | B | B | B+ | B | B- | B |
| MetaLlama (open) | B- | C+ | D | B | C+ | C+ |
| xAIGrok | C | C | B- | C+ | C- | C |
| Mistral(open) | C+ | C | D+ | C | C- | C |
| DeepSeek(open) | D+ | D | D | C- | D | D+ |
Key: A = strong, verifiable controls · B = solid, some gaps · C = partial / limited disclosure · D = minimal published control. Open-weight developers are capped on weight security by design.
How we grade
- Safety framework — a published, capability-triggered policy (RSP / Preparedness / FSF), and whether it has ever been activated.
- Capability testing — internal dangerous-capability evals plus independent pre-deployment testing (METR, Apollo, government institutes).
- Weight security — closed/structured access (revocable, monitorable) vs open weights (no post-release control); hardening toward RAND security levels.
- Transparency — system cards, safety cases, and honest capability/risk disclosures per model.
- Oversight — external audit participation, whistleblower protections, researcher safe-harbor posture.
The read on each tier
Top tier — Anthropic, OpenAI, Google DeepMind. All three publish activated capability frameworks, run internal and third-party evals, keep weights closed, and ship detailed system cards. Anthropic edges ahead on having activated a higher safety tier (ASL-3, May 2025) precautionarily and on whistleblower/transparency posture; OpenAI and DeepMind are close behind with mature Preparedness and Frontier Safety frameworks.
Middle — Microsoft. As a deployer of frontier models (and its own), it inherits strong closed-access controls and enterprise security, but publishes less standalone frontier-safety framework detail than the three labs above.
Open-weight tier — Meta, Mistral, DeepSeek. These score lower on control for a structural reason: once weights are released, revocation, monitoring, and shutdown no longer apply, and safety fine-tuning can be stripped. Meta rates highest here for responsible-use documentation and some evals; DeepSeek lowest for minimal published safety framework or third-party testing. This is a security-posture judgment, not a verdict on model quality — open weights carry real benefits for transparency and competition.
xAI sits mid-low: closed access helps its weight-security grade, but its published safety framework and third-party-testing participation are less developed than the top tier.
FAQ
Which AI model is the most secure?
By published control and transparency, Anthropic’s Claude rates highest in our 2026 scorecard, followed closely by OpenAI’s GPT and Google DeepMind’s Gemini — all closed-access models with activated safety frameworks and third-party testing.
Are open-source AI models less secure?
From a control standpoint, yes — once weights are public, they can’t be recalled, monitored, or shut off, and safety guardrails can be removed. That’s a deliberate trade-off for openness, not necessarily a sign of a worse model.
Is this the same as a best AI model ranking?
No. This rates security posture and control, not capability. A model can be highly capable and still score lower here if its developer discloses less or releases open weights.
Sources
OpenAI Preparedness Framework — openai.com
Google DeepMind Frontier Safety Framework — deepmind.google
AI Seoul Summit, Frontier AI Safety Commitments — gov.uk
RAND, Securing AI Model Weights — rand.org
METR model evaluations — metr.org
California SB 53 — leginfo.legislature.ca.gov