AI Year in Review · Oct 10, 2025 – Oct 10, 2026

Models & LLM releases

122 stories in this category.

Claude Haiku 5.5 (Anthropic)

Claude Haiku 5.5 at 10 cents per million input tokens

GLiDE (Fastino)

Fastino GLiDE, a thinking decision model

Mistral Large 4 (Mistral AI)

Mistral Large 4 "Le Chonk": 1T parameters, open weights promised for October

Beam (Reflection AI)

Reflection AI announces Beam, a 501B Western open-weight model

Rho-1 (Reka)

Reka Rho-1, a 19B research-preview omni model

GPT-6 in ChatGPT for everyone (OpenAI)

GPT-6 Sol (paid) & GPT-6 Luna (Free/Go) roll out to all ChatGPT users with "Intelligent UI".

Claude Sonnet 5.5 (Anthropic)

Anthropic ships Claude Sonnet 5.5, beating Opus 5.5 on Terminal-Bench 4.0

Gemini 4 Argon (Google DeepMind)

Google's Gemini 4 Argon tops Text Arena and the Vals Index, trusted testers only

D1 (Liquid AI)

Liquid AI releases D1, its first decision model

GPT-6.1 Sol (OpenAI)

OpenAI releases GPT-6.1 Sol: near-Astra intelligence, cached input 95% off

Perceptron Mk1.5 (Perceptron)

Perceptron Mk1.5 arrives on OpenRouter

Span-01 (Respan)

Respan's Span-01 decision model claims half the price of Jev

Gemini 4 Argon

Google released Gemini 4 Argon first to trusted cyber defenders, with a 1-million-token output limit . The same day, Anthropic’s release notes deprecated Claude Sonnet 4.5, with retirement on November 30, 2026.

GPT-6.1 Sol

OpenAI upgraded Sol a week after its release.

OpenAI DevDay

OpenAI’s DevDay recap covers GPT-6.1 Sol, always-on Dots agents and an Agents API with Computer Use .

Claude Sonnet 5.5

Anthropic released Sonnet 5.5, the first Sonnet model to launch with cyber safeguards .

Meta Connect

Meta said its Muse agent is coming to its AI glasses “in the coming months.”

Claude Opus 5.5

Anthropic released Opus 5.5 with a 1-million-token context window . Artificial Analysis scored it 58 on its Intelligence Index at max effort.

GPT-6 Sol and GPT-6 Luna

OpenAI added two GPT-6 models in ChatGPT Work, Codex and the API .

xAI Grok 4.7

xAI released Grok 4.7 with a 500,000-token context window and reported 71.0% on DeepSWE v1.1 at high effort in its own testing.

Union Alpha (OpenRouter)

Union Alpha: anonymous stealth model free on OpenRouter with 262K context

Jev (TypeSafe AI)

TypeSafe AI launches Jev, the first public non-LLM System 1 decision model

SWE-2 (Cognition)

Cognition SWE-2: near-frontier coding at up to 70% lower cost, free for a month on Devin

DeepSeek-V4.1-Flash

DeepSeek’s change log calls it “the smallest model in our new architecture family, with native multimodal visual understanding”; it replaces V4 Flash and V4 Flash Vision Exp.

Qwen3.8-Max-0902 (Alibaba Qwen)

Qwen3.8-Max-0902: 2.4T API-only refresh claims #1 on Code Arena

Claude Fable 5.1 & Mythos 5.1 (Anthropic)

Claude Fable 5.1 & Mythos 5.1: Terminal-Bench 4.0 jumps to 55.8, cache reads 75% cheaper

Gemini 3.8 Flash (Google DeepMind)

Gemini 3.8 Flash: third Flash in three weeks, HLE-Verified 54.9, 1M context

Muse Spark 1.3 (Meta AI)

Meta Muse Spark 1.3 ties GPT-5.6 Sol and Grok 4.6 on the AA index, max mode ties Fable 5

GPT-6 Astra (OpenAI)

OpenAI launches GPT-6 Astra, calls it its most intelligent and aligned model

Solaris (Runway)

Runway Solaris: an Interface World Model that generates clickable UIs frame by frame

Gemini 3.8 Flash and 3.8 Flash Cyber

Google released the pair: Flash across its developer tools and consumer products, Flash Cyber only for trusted defenders through its Fairwind Program.

Meta Muse Spark 1.3

Meta announced Muse Spark 1.3, available on Muse Code and the Meta Model API.

Claude Fable 5.1 and Claude Mythos 5.1

Anthropic released both, saying they “are the same model, but with different levels of safeguards.” Fable 5.1 is generally available; Mythos 5.1 is limited to vetted organizations.

Navigator n2 (Yutori)

Yutori Navigator n2: 27B computer-use model at a fraction of frontier cost

OpenAI speech-to-text deprecations

OpenAI’s deprecations page lists whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize as deprecated, with shutdown on February 26, 2027. The recommended replacements are gpt-li

DeepSeek-V4-Flash-Vision-Exp

The same change log adds DeepSeek-V4-Flash-Vision-Exp as an experimental model in the API.

Gemini 3.7 Flash (Google DeepMind)

Gemini 3.7 Flash: 300+ tok/s mid-tier multimodal with a 50% price cut

Grok 4.6 (xAI)

Grok 4.6 ties GPT 5.6 Sol at half the price

Gemini 3.7 Flash

Google’s Gemini API changelog lists Gemini 3.7 Flash as generally available in the Gemini API, with a 1,048,576-token context window and text, image, video, audio and PDF input.

DeepSeek-V4-Pro

DeepSeek’s API change log records DeepSeek-V4-Pro as generally available across the DeepSeek app, web and API.

Claude Opus 4.1 retired

Anthropic’s release notes list Claude Opus 4.1 as retired: requests to the model now return an error, and Anthropic recommends Claude Opus 5 as the replacement.

Qwen3.8-Max (Alibaba (Qwen))

Qwen3.8-Max: Alibaba's 2.4T-parameter flagship, with open weights promised within a week

DeepSeek-V4-Flash

DeepSeek’s API change log records DeepSeek-V4-Flash as a public beta in the DeepSeek API.

Claude Opus 5 (Anthropic)

Anthropic ships Claude Opus 5 at unchanged Opus pricing, claiming near-Fable-5 performance at half the cost

Ling-3.0-Flash (Ant Group)

Ant Group's Ling-3.0-Flash: a 124B MoE that claims to match its 1T flagship on ~5B active parameters

Gemini 3.6 Flash / 3.5 Flash-Lite / 3.5 Flash Cyber (Google DeepMind)

Google ships a three-model Gemini Flash refresh — but still no Gemini 3.5 Pro

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite

Google’s Gemini API changelog lists both models as generally available in the Gemini API, each with a 1,048,576-token context window and text, image, video, audio and PDF input.

Qwen3.8-Max Preview (Alibaba Qwen)

Alibaba previews Qwen3.8-Max at a claimed 2.4T parameters, 'second only to Fable 5'

Moonshot AI Kimi K3

Moonshot introduced Kimi K3, a 2.8-trillion-parameter mixture-of-experts model with a 1-million-token context window and native image and video input. The weights are published on Hugging Face under the Kimi K3 License.

Muse Spark 1.1 & Meta Model API (Meta AI)

Meta launches Muse Spark 1.1 and its first paid Meta Model API

GPT-5.6 (Sol, Terra, Luna) (OpenAI)

OpenAI launches GPT-5.6 publicly as three tiers: Sol, Terra and Luna

OpenAI GPT-5.6

OpenAI released GPT-5.6 in three tiers: Sol, which it calls “our flagship,” Terra, “a balanced model for everyday work,” and Luna, which the API changelog positions “for efficient, high-volume workloads.” All three are in the Ope

SWE-1.7 (Cognition)

Cognition ships SWE-1.7 at 1000 tokens per second

Grok 4.5 (xAI)

SpaceXAI launches Grok 4.5, a coding-and-agents model trained with Cursor

Sonnet 5 (Anthropic)

Claude Sonnet 5: 'our most agentic Sonnet yet' at intro pricing

Claude Sonnet 5 (Anthropic)

Most agentic Sonnet, near Opus 4.8 at $2/$10; default for Free & Pro.

Claude Fable 5 & Mythos 5 (Anthropic)

Anthropic’s most capable GA model ($10/$50), plus Mythos tier above Opus with 1M context.

MAI-Code-1-Flash (Microsoft)

Microsoft ships MAI-Code-1-Flash into GitHub Copilot

MAI-Thinking-1 (Microsoft)

Microsoft launches MAI-Thinking-1, a 1T MoE trained from scratch

MiniMax M3 (MiniMax)

MiniMax announces M3 coding/agentic model with 1M context

Claude Opus 4.8 (Anthropic)

Anthropic ships Claude Opus 4.8 live mid-show

Qwen 3.7-Max (Alibaba (Qwen))

Alibaba releases Qwen 3.7-Max agentic frontier model with robotics demos

Composer 2.5 (Cursor)

Cursor launches Composer 2.5 with Opus-class coding at much lower cost

Gemini 3.5 Flash (Google DeepMind)

Gemini 3.5 Flash launches at I/O as Google's agentic workhorse model

Google I/O 2026: Gemini 3.5 Flash, Gemini Omni, Spark

"Agentic Gemini era": Gemini 3.5, Omni (reason+create), Search information agents, Antigravity 2.0.

ERNIE 5.1 Preview (Baidu)

Baidu ERNIE 5.1 Preview hits #13 on Arena with 6% of the compute

GPT-5.5 (OpenAI)

GPT-5.5 and GPT-5.5 Pro drop live, SOTA across the board

GPT-5.5 & GPT-5.5 Pro (OpenAI)

82.7% Terminal-Bench 2.0, 58.6% SWE-Bench Pro; API on Apr 24.

Claude Opus 4.7 (Anthropic)

Claude Opus 4.7 drops live with 87.6% SWE-bench Verified and xhigh effort

Muse Spark (Meta (Meta Superintelligence Labs))

Meta launches Muse Spark, first model from Meta Superintelligence Labs

Qwen3.6-Plus (Alibaba (Qwen))

Alibaba ships Qwen3.6-Plus with near-Opus agentic coding and 1M context

Composer 2 (Cursor)

Cursor Composer 2 beats Opus 4.6 on TerminalBench at a tenth of the price

MiniMax M2.7 (MiniMax)

MiniMax M2.7: first self-evolving model hits 56% on SWE-Bench Pro

GPT-5.4 Mini & Nano (OpenAI)

OpenAI ships GPT-5.4 Mini and Nano for coding, computer use, and subagents

MiMo (Xiaomi)

Xiaomi MiMo revealed as the 1T-param stealth model topping OpenRouter

GPT-5.4 mini & nano (OpenAI)

Cheaper, faster GPT-5.4 variants.

Gemini Embedding 2 (Google)

Google launches Gemini Embedding 2, a natively multimodal embedder

embed-large-v3 (Mixbread)

Mixbread embed-large-v3 beats Gemini Embedding 2

SWE-1.6 (Cognition)

Cognition previews SWE-1.6, hitting 51% on SWE Bench Pro

Gemini 3.1 Flash-Lite (Google DeepMind)

Google launches Gemini 3.1 Flash-Lite with 1M context at 360 tok/s

GPT-5.3 Instant (OpenAI)

OpenAI rolls out GPT-5.3 Instant as the free-tier fast model

GPT-5.4 (OpenAI)

OpenAI drops GPT-5.4 Thinking and GPT-5.4 Pro live during the show

Claude Sonnet 4.6 (Anthropic)

Anthropic ships Claude Sonnet 4.6 with 79.6% SWE-Bench and 1M context

Seed 2.0 (ByteDance)

ByteDance Seed 2.0: frontier multimodal family at 73-84% lower pricing

Gemini 3.1 Pro (Google DeepMind)

Gemini 3.1 Pro drops live with 44% HLE and 77% ARC-AGI at the same price

Grok 4.20 (xAI)

xAI silently drops Grok 4.20 with four 500B-param collaborating agents

Gemini 3.1 Pro (Google)

Reasoning upgrade, 77.1% ARC-AGI-2.

Claude Sonnet 4.6 (Anthropic)

Sonnet upgrade with 1M context (beta), coding & computer use.

Gemini 3 Deep Think (Google DeepMind)

Gemini 3 Deep Think scores 84% on ARC-AGI 2

GPT 5.3 Codex Spark (OpenAI)

OpenAI ships GPT 5.3 Codex Spark on Cerebras for real-time coding

Claude Opus 4.6 (Anthropic)

Anthropic ships Claude Opus 4.6 with 1M context and agent teams

GPT-5.3-Codex (OpenAI)

OpenAI answers Opus with GPT-5.3-Codex, first model that helped build itself

Gemini 2.5 (Google DeepMind)

Gemini 2.5 takes the #1 benchmark spot in March

Gemini 3 Flash (Google DeepMind)

Gemini 3 Flash delivers frontier intelligence at $0.50/1M input tokens

Mistral OCR 3 (Mistral AI)

Mistral OCR 3 claims 74% win-rate over OCR v2 with aggressive pricing

GPT 5.2 Codex (OpenAI)

GPT 5.2 Codex drops live during the show with 400K context

Gemini 3 Flash (Google)

Fast/cheap tier with strong baseline.

GPT-5.2 (OpenAI)

Instant/Thinking/Pro after "code red"; 100% AIME 2025, 70.9% GDPval.

Claude Opus 4.5 (Anthropic)

Anthropic launches Claude Opus 4.5, reclaiming the coding crown

GPT-5.1 (OpenAI)

OpenAI launches GPT-5.1 with a warmer, more personable voice

SWE-1.5 (Cognition)

Cognition SWE-1.5: 950 tok/s coding model hitting 40% on SWE-bench Pro

Claude Haiku 4.5 (Anthropic)

Claude Haiku 4.5: fast, cheap model rivals Sonnet 4 accuracy

SWE-grep (Cognition)

Cognition SWE-grep: RL-trained fast context retrieval for coding agents

OpenPipe Qwen3 14B Instruct (OpenPipe (Weights & Biases))

OpenPipe Qwen3 14B Instruct lands on W&B Inference

RTFM (World Labs)

World Labs RTFM renders 3D worlds in real time on a single H100