Claude Haiku 5.5 (Anthropic)
Claude Haiku 5.5 at 10 cents per million input tokens
AI Year in Review · Oct 10, 2025 – Oct 10, 2026
122 stories in this category.
Claude Haiku 5.5 at 10 cents per million input tokens
Fastino GLiDE, a thinking decision model
Mistral Large 4 "Le Chonk": 1T parameters, open weights promised for October
Reflection AI announces Beam, a 501B Western open-weight model
Reka Rho-1, a 19B research-preview omni model
GPT-6 Sol (paid) & GPT-6 Luna (Free/Go) roll out to all ChatGPT users with "Intelligent UI".
Anthropic ships Claude Sonnet 5.5, beating Opus 5.5 on Terminal-Bench 4.0
Google's Gemini 4 Argon tops Text Arena and the Vals Index, trusted testers only
Liquid AI releases D1, its first decision model
OpenAI releases GPT-6.1 Sol: near-Astra intelligence, cached input 95% off
Perceptron Mk1.5 arrives on OpenRouter
Respan's Span-01 decision model claims half the price of Jev
Google released Gemini 4 Argon first to trusted cyber defenders, with a 1-million-token output limit . The same day, Anthropic’s release notes deprecated Claude Sonnet 4.5, with retirement on November 30, 2026.
OpenAI upgraded Sol a week after its release.
OpenAI’s DevDay recap covers GPT-6.1 Sol, always-on Dots agents and an Agents API with Computer Use .
Anthropic released Sonnet 5.5, the first Sonnet model to launch with cyber safeguards .
Qwen 3.8 Omni Flash with 1M context and Qwen 3.8 Live Translate
Anthropic launches Claude Opus 5.5: Fable-level work for 40% less
OpenAI releases GPT-6 Luna at $0.10 / $0.50 per million tokens
OpenAI rolls out GPT-6 Sol at half the GPT-5.6 price
xAI releases Grok 4.7 with 46% on CursorBench
Meta said its Muse agent is coming to its AI glasses “in the coming months.”
Anthropic released Opus 5.5 with a 1-million-token context window . Artificial Analysis scored it 58 on its Intelligence Index at max effort.
OpenAI added two GPT-6 models in ChatGPT Work, Codex and the API .
xAI released Grok 4.7 with a 500,000-token context window and reported 71.0% on DeepSWE v1.1 at high effort in its own testing.
Union Alpha: anonymous stealth model free on OpenRouter with 262K context
TypeSafe AI launches Jev, the first public non-LLM System 1 decision model
Cognition SWE-2: near-frontier coding at up to 70% lower cost, free for a month on Devin
DeepSeek’s change log calls it “the smallest model in our new architecture family, with native multimodal visual understanding”; it replaces V4 Flash and V4 Flash Vision Exp.
Qwen3.8-Max-0902: 2.4T API-only refresh claims #1 on Code Arena
Claude Fable 5.1 & Mythos 5.1: Terminal-Bench 4.0 jumps to 55.8, cache reads 75% cheaper
Gemini 3.8 Flash: third Flash in three weeks, HLE-Verified 54.9, 1M context
Meta Muse Spark 1.3 ties GPT-5.6 Sol and Grok 4.6 on the AA index, max mode ties Fable 5
OpenAI launches GPT-6 Astra, calls it its most intelligent and aligned model
Runway Solaris: an Interface World Model that generates clickable UIs frame by frame
Google released the pair: Flash across its developer tools and consumer products, Flash Cyber only for trusted defenders through its Fairwind Program.
Meta announced Muse Spark 1.3, available on Muse Code and the Meta Model API.
Anthropic released both, saying they “are the same model, but with different levels of safeguards.” Fable 5.1 is generally available; Mythos 5.1 is limited to vetted organizations.
Yutori Navigator n2: 27B computer-use model at a fraction of frontier cost
OpenAI’s deprecations page lists whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize as deprecated, with shutdown on February 26, 2027. The recommended replacements are gpt-li
The same change log adds DeepSeek-V4-Flash-Vision-Exp as an experimental model in the API.
Gemini 3.7 Flash: 300+ tok/s mid-tier multimodal with a 50% price cut
Grok 4.6 ties GPT 5.6 Sol at half the price
Google’s Gemini API changelog lists Gemini 3.7 Flash as generally available in the Gemini API, with a 1,048,576-token context window and text, image, video, audio and PDF input.
DeepSeek’s API change log records DeepSeek-V4-Pro as generally available across the DeepSeek app, web and API.
Anthropic’s release notes list Claude Opus 4.1 as retired: requests to the model now return an error, and Anthropic recommends Claude Opus 5 as the replacement.
Qwen3.8-Max: Alibaba's 2.4T-parameter flagship, with open weights promised within a week
DeepSeek’s API change log records DeepSeek-V4-Flash as a public beta in the DeepSeek API.
Anthropic ships Claude Opus 5 at unchanged Opus pricing, claiming near-Fable-5 performance at half the cost
Ant Group's Ling-3.0-Flash: a 124B MoE that claims to match its 1T flagship on ~5B active parameters
Google ships a three-model Gemini Flash refresh — but still no Gemini 3.5 Pro
Google’s Gemini API changelog lists both models as generally available in the Gemini API, each with a 1,048,576-token context window and text, image, video, audio and PDF input.
Alibaba previews Qwen3.8-Max at a claimed 2.4T parameters, 'second only to Fable 5'
Moonshot introduced Kimi K3, a 2.8-trillion-parameter mixture-of-experts model with a 1-million-token context window and native image and video input. The weights are published on Hugging Face under the Kimi K3 License.
Meta launches Muse Spark 1.1 and its first paid Meta Model API
OpenAI launches GPT-5.6 publicly as three tiers: Sol, Terra and Luna
OpenAI released GPT-5.6 in three tiers: Sol, which it calls “our flagship,” Terra, “a balanced model for everyday work,” and Luna, which the API changelog positions “for efficient, high-volume workloads.” All three are in the Ope
Cognition ships SWE-1.7 at 1000 tokens per second
SpaceXAI launches Grok 4.5, a coding-and-agents model trained with Cursor
Base44 launches Base 1, the first in-house vibe-coding LLM
OpenAI ships GPT-5.6 as a three-model family: Sol, Terra and Luna
Fable 5 restored globally after the export-control pause
Claude Sonnet 5: 'our most agentic Sonnet yet' at intro pricing
Most agentic Sonnet, near Opus 4.8 at $2/$10; default for Free & Pro.
Anthropic’s most capable GA model ($10/$50), plus Mythos tier above Opus with 1M context.
Microsoft ships MAI-Code-1-Flash into GitHub Copilot
Microsoft launches MAI-Thinking-1, a 1T MoE trained from scratch
MiniMax announces M3 coding/agentic model with 1M context
Anthropic ships Claude Opus 4.8 live mid-show
Alibaba releases Qwen 3.7-Max agentic frontier model with robotics demos
Cursor launches Composer 2.5 with Opus-class coding at much lower cost
Gemini 3.5 Flash launches at I/O as Google's agentic workhorse model
"Agentic Gemini era": Gemini 3.5, Omni (reason+create), Search information agents, Antigravity 2.0.
Baidu ERNIE 5.1 Preview hits #13 on Arena with 6% of the compute
Mayo Clinic's REDMOD detects pancreatic cancer 3 years early
OpenAI releases clinician/medical model and workspace agents
GPT-5.5 and GPT-5.5 Pro drop live, SOTA across the board
82.7% Terminal-Bench 2.0, 58.6% SWE-Bench Pro; API on Apr 24.
Claude Opus 4.7 drops live with 87.6% SWE-bench Verified and xhigh effort
Meta launches Muse Spark, first model from Meta Superintelligence Labs
Alibaba ships Qwen3.6-Plus with near-Opus agentic coding and 1M context
Cursor Composer 2 beats Opus 4.6 on TerminalBench at a tenth of the price
MiniMax M2.7: first self-evolving model hits 56% on SWE-Bench Pro
OpenAI ships GPT-5.4 Mini and Nano for coding, computer use, and subagents
Xiaomi MiMo revealed as the 1T-param stealth model topping OpenRouter
Cheaper, faster GPT-5.4 variants.
Google launches Gemini Embedding 2, a natively multimodal embedder
Mixbread embed-large-v3 beats Gemini Embedding 2
Cognition previews SWE-1.6, hitting 51% on SWE Bench Pro
Google launches Gemini 3.1 Flash-Lite with 1M context at 360 tok/s
OpenAI rolls out GPT-5.3 Instant as the free-tier fast model
OpenAI drops GPT-5.4 Thinking and GPT-5.4 Pro live during the show
Anthropic ships Claude Sonnet 4.6 with 79.6% SWE-Bench and 1M context
ByteDance Seed 2.0: frontier multimodal family at 73-84% lower pricing
Gemini 3.1 Pro drops live with 44% HLE and 77% ARC-AGI at the same price
xAI silently drops Grok 4.20 with four 500B-param collaborating agents
Reasoning upgrade, 77.1% ARC-AGI-2.
Sonnet upgrade with 1M context (beta), coding & computer use.
Gemini 3 Deep Think scores 84% on ARC-AGI 2
OpenAI ships GPT 5.3 Codex Spark on Cerebras for real-time coding
Anthropic ships Claude Opus 4.6 with 1M context and agent teams
OpenAI answers Opus with GPT-5.3-Codex, first model that helped build itself
Claude Opus 4 drops in Q2 — Ryan's pick for best model ever
Gemini 2.5 takes the #1 benchmark spot in March
MiniMax drops Hailuo 2.3 in November
GPT-5 Codex: OpenAI's specialized coding model moves the stock
Gemini 3 Flash delivers frontier intelligence at $0.50/1M input tokens
Mistral OCR 3 claims 74% win-rate over OCR v2 with aggressive pricing
GPT 5.2 Codex drops live during the show with 400K context
Fast/cheap tier with strong baseline.
Instant/Thinking/Pro after "code red"; 100% AIME 2025, 70.9% GDPval.
Anthropic launches Claude Opus 4.5, reclaiming the coding crown
Gemini 3 Pro launches with record ARC-AGI-2 scores
GPT-5.1-Codex-Max runs 24-hour coding tasks with native compaction
Sunday Robotics unveils ACT-1 home robot foundation model and Memo
Grok 4.1 briefly tops LM Arena with major post-training upgrade
OpenAI launches GPT-5.1 with a warmer, more personable voice
Cognition SWE-1.5: 950 tok/s coding model hitting 40% on SWE-bench Pro
Claude Haiku 4.5: fast, cheap model rivals Sonnet 4 accuracy
Cognition SWE-grep: RL-trained fast context retrieval for coding agents
OpenPipe Qwen3 14B Instruct lands on W&B Inference
World Labs RTFM renders 3D worlds in real time on a single H100