Foundation Models
Model releases, benchmark results, pricing changes, and open-weight developments. Curated for builders who need to track what is available and what it costs.
OpenAI Launches GPT-5.6 Sol, Terra, Luna - Limited to Trusted Partners at US Government Request
OpenAI announces three new frontier models - GPT-5.6 Sol (strongest model yet, with top performance on coding and cybersecurity), Terra, and Luna - but limits the launch to a small group of trusted partners at the US government's request. OpenAI previewed capabilities with the government ahead of launch and says it is developing a repeatable process for future model releases under this framework.
Why it matters: Government-gated frontier model launches are becoming a structural feature of the AI landscape. Following the brief suspension of Claude Fable 5 weeks earlier, GPT-5.6's restricted launch confirms the US is establishing a national security review process for the most capable AI systems. For enterprises and developers, this means top-tier model access will increasingly depend on trust tier and jurisdiction - not just API keys.
Claude Fable 5 - Anthropic's Frontier Model Launched Then Briefly Pulled by US Export Order
Anthropic releases Claude Fable 5 - state-of-the-art across coding, scientific research, vision, and knowledge work at $10/$50 per 1M tokens (less than half Claude Mythos Preview pricing). Within three days of launch, a US government export directive temporarily forced it offline; it was reinstated June 22 for Pro, Max, Team, and Enterprise subscribers.
Why it matters: Fable 5 sets a new capability ceiling for commercially available models. The temporary government-directed suspension - the first of its kind for a major AI model - signals the US is moving toward active national security oversight of frontier model deployments, not just policy guidance. This precedent will shape how Anthropic and other labs release their most capable models going forward.
Microsoft Build 2026 - MAI-Thinking-1 and MAI-Code-1-Flash, First In-House Models
Microsoft announces 7 MAI models at Build 2026. MAI-Thinking-1 (35B active params, 256K context) is Microsoft's first in-house reasoning model - trained without OpenAI distillation, preferred over Claude Sonnet 4.6 in blind evals. MAI-Code-1-Flash (5B params) integrates deep into GitHub Copilot and VS Code for all plans.
Why it matters: Microsoft's first home-built frontier-class models mark a strategic shift away from pure OpenAI dependency. MAI-Code-1-Flash becoming the backbone of GitHub Copilot affects every developer using Copilot - Microsoft controls the model and pricing, not OpenAI.
Claude Opus 4.8 - #1 on AI Intelligence Index, Dynamic Workflows for Claude Code
Anthropic releases Claude Opus 4.8 at $5/$25 per 1M tokens - the only model to complete every case end-to-end on the Super-Agent benchmark. Scores 84% on Online-Mind2Web (browser agent eval). Introduces Dynamic Workflows in Claude Code for very large-scale tasks. Fast mode now 3ร cheaper than for previous Opus.
Why it matters: Opus 4.8 is the new benchmark leader for agentic and computer-use tasks. For builders running autonomous coding agents or browser agents, it outperforms GPT-5.5 on key evals while matching it on cost. The 4ร reduction in unremarked code flaws vs Opus 4.7 matters for production code review workflows.
Gemini 3.5 Flash GA - Google's Speed/Quality Leader for Agents and Coding
Gemini 3.5 Flash reaches general availability following its Google I/O announcement. Achieves 284 tokens/second and an Intelligence Index score of 55 - strong coding and agentic performance at Flash-tier pricing. Gemini 3.5 Pro (2M token context, Deep Think reasoning mode) committed for June GA.
Why it matters: Gemini 3.5 Flash is now the natural default for Google-stack builders who need fast agentic throughput. The 2M context window on Pro will make whole-codebase and large document workflows more practical than any current model.
Qwen3 - Alibaba's New Model Family with Thinking Mode
Alibaba releases Qwen3 family: 0.6B to 235B MoE. All models support a 'thinking mode' toggle (chain-of-thought on/off per request). Qwen3-Coder 480B MoE targets software engineering. Apache 2.0 licensed. Strong in 29 languages.
Why it matters: Thinking-mode toggle is a practical innovation - use fast mode for simple tasks, reasoning mode for complex ones, in the same model. Qwen3 Coder 480B directly challenges frontier coding models at open-weight prices. Extends Chinese AI labs' impact on the open-weight ecosystem.
Grok 4.3 - xAI's Updated Flagship with 1M Context
xAI releases Grok 4.3 - current flagship at $1.25/$2.50 per 1M tokens with a 1M token context window. Positioned between Claude Sonnet and Opus on price/quality. $150/month in free developer credits available via data-sharing program.
Why it matters: 1M context window at competitive pricing makes Grok 4.3 a viable choice for long-document tasks. The free $150/month credit program is the most generous developer offer from any major AI lab in 2026 - lowering the barrier to experiment with xAI's API significantly.
GPT-5.5 Launches - OpenAI's Most Capable General Model, $5/$30 per 1M Tokens
GPT-5.5 launches across ChatGPT and API, priced at $5/$30 per 1M tokens. It matches GPT-5.4 latency while performing at a significantly higher level - excels at multi-step agentic tasks, coding, and tool use, using fewer tokens per Codex task than its predecessor.
Why it matters: Sets a new capability/efficiency bar for paid frontier models. The combination of higher intelligence and lower token usage makes it more economical for agentic pipelines than it first appears from headline pricing. GPT-5.5 Pro ($30/$180) targets research-grade complexity.
Llama 4 Scout & Maverick - Meta Ships Open-Weight Multimodal MoE
Meta releases Llama 4 Scout (109B MoE, 10M token context) and Maverick (400B MoE, 17B active parameters). Both are natively multimodal, Apache 2.0 licensed, and match or beat GPT-4o on major benchmarks at a fraction of the inference cost.
Why it matters: First open-weight models to seriously challenge frontier closed-source models on quality. The 10M context window on Scout is the largest of any openly available model. MoE architecture means inference cost scales with active parameters (17B), not total (400B). Shifts the open vs closed model debate significantly.
DeepSeek Releases R2 - Open-Weight Reasoning Model
DeepSeek R2 achieves competitive reasoning performance with an open-weight license, making advanced reasoning accessible to self-hosted deployments.
Why it matters: Open-weight reasoning models reduce dependency on closed APIs for complex tasks. Important for enterprises with data residency requirements.
GPT-5 Launches - OpenAI Frontier Model with 400K Token Context
GPT-5 launches as OpenAI's new flagship with a 400K token context window, strong AIME 2025 maths performance, and significantly improved multi-step project execution and autonomous coding capability.
Why it matters: Sets a new capability baseline for closed frontier models. The 400K context window makes whole-codebase and large document reasoning practical via API. Forces pricing and capability recalibration across all competing providers.
Anthropic Releases Claude Opus 4 - Most Capable Model Yet
Claude Opus 4 sets new benchmarks across coding, reasoning, and extended thinking tasks, with improved tool use and agentic capabilities.
Why it matters: Represents a significant step in model capability for builders relying on agentic workflows and complex multi-step reasoning.
OpenAI Releases o3 and o4-mini - Reasoning Models with Native Tool Use
o3 and o4-mini combine chain-of-thought reasoning with native tool use, enabling models to search the web, run code, and call APIs mid-reasoning.
Why it matters: Reasoning + tool use in a single model removes the need to orchestrate separate search and reasoning steps, simplifying agentic pipeline design.
Meta Releases Llama 4 - Natively Multimodal Open-Weight MoE Models
Meta releases Llama 4 Scout (17B active params, 10M token context, runs on a single H100) and Maverick (17B active/400B total, 1M context) - the first natively multimodal Llama models trained on text, images, and video data.
Why it matters: Llama 4 is the new open-weight baseline for self-hosted multimodal deployments. Enterprises with data residency requirements now have a competitive open alternative to closed frontier models at a fraction of the API cost.
Google Gemini 2.5 Pro - 1M Token Context and Thinking Mode Released
Gemini 2.5 Pro adds a thinking mode (extended reasoning) alongside its 1M token context window, topping key benchmarks including coding and maths.
Why it matters: 1M context makes whole-codebase and whole-document analysis practical. Thinking mode brings reasoning capability to Google's ecosystem.