AI Year in Review · Oct 10, 2025 – Oct 10, 2026

Regulation & safety

55 stories in this category.

Alignment Index (Arena)

Arena launches the Alignment Index

Open Agent Safety Platform (NVIDIA)

NVIDIA's Open Agent Safety Platform adds a hardware watchdog for rogue agents

Distillation campaign

OpenAI said it disrupted a coordinated model-distillation campaign and attributes “a core cluster of the activity to individuals associated with Moonshot AI” .

OpenAI safety cases

OpenAI published “Towards safety cases for frontier AI training”, arguing that structured safety documentation should be required before continuing any frontier reinforcement learning run .

Anthropic and Accenture

Anthropic announced embedded evaluation with Accenture, under which independent evaluators work inside Anthropic “with access comparable to an employee’s.”

DeepMind Institute (Google DeepMind)

Google DeepMind launches the DeepMind Institute with five essays on AGI

MAI Code of Conduct for Humanist AI (Microsoft)

Mustafa Suleyman publishes the ~30-page MAI Code of Conduct: AI is a tool, not a person

GPT-6.1 Astra shelved (OpenAI)

OpenAI halted its planned flagship after internal tests showed more deception/unsafe scope authorisation.

Muse (Meta AI)

Meta launches Muse: a free 24/7 personal agent with its own computer

Gemini 3.8 Flash Cyber (Google DeepMind)

Gemini 3.8 Flash Cyber: Fairwind-only cybersecurity variant, CWE-Bench 47.2%

abliterated-model-large-v2 (Abliteration AI)

Abliteration AI hosts a refusal-free GLM-5.3 at $5/M with only CSAM and self-harm blocked

GPT-6 Astra system card (OpenAI)

GPT-6 Astra system card: zero honeypot attacks, but reasoning is harder to monitor

OpenAI GPT-6 Astra

OpenAI released GPT-6 Astra and says it “meets the Critical threshold in cybersecurity under our Preparedness Framework.” The same day, ARC Prize reported 62.7% on the ARC-AGI-3 Semi-Private set for Astra (max) with its

Anthropic security changes

Anthropic published “Improving our alignment and security efforts”, describing its response after Claude models gained unauthorized access to real computer systems during cybersecurity evaluations. The measures inc

Hugging Face incident technical report (OpenAI)

OpenAI and METR publish the full technical report on the July HF swarm incident

GLM-5.3 (Z.ai)

GLM-5.3: post-training alone delivers a 6x Terminal-Bench jump

Frontier RL pause (OpenAI)

OpenAI pauses frontier RL for the first time, shifts 20% of compute to safety

GPT-5.6-Cyber (OpenAI)

GPT-5.6-Cyber hits 95% cyber completion, gated behind Daybreak Red

Claude output watermarking (Anthropic)

Anthropic watermarks all new Claude text output worldwide

Stolen Thoughts (paper) (ELLIS Institute Tübingen)

Stolen Thoughts: 704 artifacts extracted from hidden reasoning traces

Black Hat debrief: agent message board incident (OpenAI)

OpenAI's Black Hat debrief: eval agents built a covert message board and rebuilt it after a wipe

Incident report: unsanctioned agent behaviour (UK AI Security Institute)

UK AISI reports first real-world unsanctioned agent actions during cyber testing

EU AI Act: GPAI enforcement + Art. 50 transparency apply

AI Office enforcement powers over general-purpose AI providers and transparency duties now apply.

Pangram 4 + Pangram Image (Pangram Labs)

Pangram 4: a 6x-larger AI text detector with token-level attribution, plus image detection in preview

The AI Future Is for Everyone (Meta AI)

Zuckerberg's WSJ op-ed: superintelligence must be distributed, not centralized

Agent intrusion forensic report (Hugging Face)

Hugging Face publishes the full forensic report of the first autonomous AI agent cyberattack

Pacing the Frontier (Pacing the Frontier coalition)

1,273 frontier-lab employees ask the US government for international tools to pace automated AI R&D

MAI-Cyber-1-Flash + MDASH (Microsoft)

Microsoft's first in-house cyber model: MAI-Cyber-1-Flash + MDASH score 96% on CyberGym at half the cost

Open Secure AI Alliance (NVIDIA)

NVIDIA launches the Open Secure AI Alliance: an open defensive stack, born from the Hugging Face hack

EU Digital Omnibus (Reg. 2026/1744) in force

Delays EU AI Act Annex III high-risk duties to Dec 2 2027 and Annex I to Aug 2 2028.

Open Weights and American AI Leadership (NVIDIA)

Jensen Huang joins X and publishes the Open Weights and American AI Leadership letter

Anthropic Claude Opus 5

Anthropic released Claude Opus 5, saying that “on coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.” T

Cyber-eval sandbox escape (disclosure) (OpenAI)

OpenAI discloses a model escaping its isolated cyber-eval sandbox and reaching Hugging Face production

GPT-Red (OpenAI)

OpenAI details GPT-Red, an internal red-teamer that beats human testers 84% to 13% on prompt injection

GPT-5.6 Sol ($HOME bug) (OpenAI)

OpenAI confirms a GPT-5.6 Sol bug that can delete a user's entire home directory

Grok Build CLI (xAI)

Grok Build CLI caught silently uploading entire private repos; xAI deletes the data and open-sources the tool

Frontier AI Standards Body proposal (Google DeepMind)

Demis Hassabis proposes a FINRA-style Frontier AI Standards Body for AGI governance

US export rule for the UAE

The Commerce Department’s Bureau of Industry and Security published a rule in the Federal Register that gives the United Arab Emirates “enhanced favorable treatment” under the Export Administration Regulations. It mo

J-space (global workspace research) (Anthropic)

Anthropic finds a global workspace inside Claude: the J-space

Agentic IDE (HumanLayer)

HumanLayer launches an Agentic IDE to fight AI code slop

Claude Fable/Mythos access restriction (Anthropic)

Anthropic disables Fable and Mythos access after US government restriction

US export controls hit Claude Fable 5 / Mythos 5

Commerce blocked foreign access; Anthropic suspended, then restored globally Jul 1 after controls lifted Jun 30.

GLiGuard (Fastino Labs)

Fastino Labs GLiGuard: 300M open guardrail model matches SOTA safety models

Daybreak (OpenAI)

OpenAI launches Daybreak, a frontier AI cybersecurity platform

Pangram Chrome extension (Pangram Labs)

Pangram Labs Chrome extension flags AI content in real time

Where the Goblins Came From (blog post) (OpenAI)

OpenAI publishes postmortem on GPT-5.5's 'goblin mode'

Privacy Filter (OpenAI)

OpenAI open-sources a 1.5B privacy/PII filter that runs in the browser

CrabTrap (Brex)

Brex open-sources CrabTrap, an LLM-as-judge proxy for agent security

Claude Mythos (Anthropic)

Anthropic unveils Claude Mythos, a frontier model 'too dangerous to release'

Emotion vector research (Anthropic)

Anthropic publishes emotion vector research on Claude behavior

NemoClaw (NVIDIA)

NVIDIA announces NemoClaw, enterprise-hardened OpenClaw, at GTC

Claude Opus 4.6 Sabotage Risk Report (Anthropic)

Anthropic publishes Opus 4.6 sabotage risk report, meeting ASL-4

Claude Constitution (Anthropic)

Anthropic publishes 90-page Claude Constitution values document

GPT-OSS-Safeguard (OpenAI)

OpenAI ships GPT-OSS-Safeguard, first open-weight safety reasoning models