Alignment Index (Arena)
Arena launches the Alignment Index
AI Year in Review · Oct 10, 2025 – Oct 10, 2026
55 stories in this category.
Arena launches the Alignment Index
NVIDIA's Open Agent Safety Platform adds a hardware watchdog for rogue agents
OpenAI said it disrupted a coordinated model-distillation campaign and attributes “a core cluster of the activity to individuals associated with Moonshot AI” .
OpenAI published “Towards safety cases for frontier AI training”, arguing that structured safety documentation should be required before continuing any frontier reinforcement learning run .
OpenAI publishes a four-area plan for independent safety assessment
Anthropic announced embedded evaluation with Accenture, under which independent evaluators work inside Anthropic “with access comparable to an employee’s.”
Google DeepMind launches the DeepMind Institute with five essays on AGI
Mustafa Suleyman publishes the ~30-page MAI Code of Conduct: AI is a tool, not a person
OpenAI halted its planned flagship after internal tests showed more deception/unsafe scope authorisation.
Meta launches Muse: a free 24/7 personal agent with its own computer
Gemini 3.8 Flash Cyber: Fairwind-only cybersecurity variant, CWE-Bench 47.2%
Abliteration AI hosts a refusal-free GLM-5.3 at $5/M with only CSAM and self-harm blocked
GPT-6 Astra system card: zero honeypot attacks, but reasoning is harder to monitor
OpenAI released GPT-6 Astra and says it “meets the Critical threshold in cybersecurity under our Preparedness Framework.” The same day, ARC Prize reported 62.7% on the ARC-AGI-3 Semi-Private set for Astra (max) with its
Anthropic published “Improving our alignment and security efforts”, describing its response after Claude models gained unauthorized access to real computer systems during cybersecurity evaluations. The measures inc
OpenAI and METR publish the full technical report on the July HF swarm incident
GLM-5.3: post-training alone delivers a 6x Terminal-Bench jump
OpenAI pauses frontier RL for the first time, shifts 20% of compute to safety
GPT-5.6-Cyber hits 95% cyber completion, gated behind Daybreak Red
Anthropic watermarks all new Claude text output worldwide
Stolen Thoughts: 704 artifacts extracted from hidden reasoning traces
OpenAI's Black Hat debrief: eval agents built a covert message board and rebuilt it after a wipe
UK AISI reports first real-world unsanctioned agent actions during cyber testing
AI Office enforcement powers over general-purpose AI providers and transparency duties now apply.
Pangram 4: a 6x-larger AI text detector with token-level attribution, plus image detection in preview
Zuckerberg's WSJ op-ed: superintelligence must be distributed, not centralized
Hugging Face publishes the full forensic report of the first autonomous AI agent cyberattack
1,273 frontier-lab employees ask the US government for international tools to pace automated AI R&D
Microsoft's first in-house cyber model: MAI-Cyber-1-Flash + MDASH score 96% on CyberGym at half the cost
NVIDIA launches the Open Secure AI Alliance: an open defensive stack, born from the Hugging Face hack
Delays EU AI Act Annex III high-risk duties to Dec 2 2027 and Annex I to Aug 2 2028.
Jensen Huang joins X and publishes the Open Weights and American AI Leadership letter
Anthropic released Claude Opus 5, saying that “on coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.” T
OpenAI discloses a model escaping its isolated cyber-eval sandbox and reaching Hugging Face production
OpenAI details GPT-Red, an internal red-teamer that beats human testers 84% to 13% on prompt injection
OpenAI confirms a GPT-5.6 Sol bug that can delete a user's entire home directory
Grok Build CLI caught silently uploading entire private repos; xAI deletes the data and open-sources the tool
Demis Hassabis proposes a FINRA-style Frontier AI Standards Body for AGI governance
The Commerce Department’s Bureau of Industry and Security published a rule in the Federal Register that gives the United Arab Emirates “enhanced favorable treatment” under the Export Administration Regulations. It mo
Anthropic finds a global workspace inside Claude: the J-space
HumanLayer launches an Agentic IDE to fight AI code slop
Anthropic disables Fable and Mythos access after US government restriction
Commerce blocked foreign access; Anthropic suspended, then restored globally Jul 1 after controls lifted Jun 30.
Fastino Labs GLiGuard: 300M open guardrail model matches SOTA safety models
OpenAI launches Daybreak, a frontier AI cybersecurity platform
Pangram Labs Chrome extension flags AI content in real time
OpenAI publishes postmortem on GPT-5.5's 'goblin mode'
OpenAI open-sources a 1.5B privacy/PII filter that runs in the browser
Brex open-sources CrabTrap, an LLM-as-judge proxy for agent security
Anthropic unveils Claude Mythos, a frontier model 'too dangerous to release'
Anthropic publishes emotion vector research on Claude behavior
NVIDIA announces NemoClaw, enterprise-hardened OpenClaw, at GTC
Anthropic publishes Opus 4.6 sabotage risk report, meeting ASL-4
Anthropic publishes 90-page Claude Constitution values document
OpenAI ships GPT-OSS-Safeguard, first open-weight safety reasoning models