Models
Every AI model release, newest first.
2026
35 models- OpenAI launches GPT-6 Sol and GPT-6 LunaBuilt with Astra's methods, GPT-6 Sol offers 'Astra-level reliability' with about half the errors of GPT-5.6 Sol, while Luna targets cheap, high-volume work; both cost half as much as the 5.6 series.
- Anthropic releases Claude Opus 5.5Claude Opus 5.5 matches Fable 5.1 on most work with a 1M-token context, runs 30% faster, cuts API prices 20% to $4/$20 per million tokens, and scored 66.4% on Terminal-Bench 4.0 versus 57.9% for GPT-6 Astra.
- SpaceXAI releases Grok 4.7Grok 4.7 moved to a larger ~2.1T-parameter base model with longer reinforcement learning on multi-hour tasks and a new safeguard stack, at the same $2/$6 pricing as Grok 4.6.
- Alibaba releases Qwen3.8-Omni-FlashQwen3.8-Omni-Flash is an API-only omni-modal model with a 1M-token context, native audio-video reasoning and tool use, aimed at agentic applications.
- DeepSeek releases V4.1-Flash with native visionDeepSeek-V4.1-Flash, the smallest model in its new V4.1 architecture family, added native multimodal visual understanding with faster inference and higher throughput, available through the DeepSeek API.
- OpenAI launches GPT-6 AstraOpenAI's GPT-6 Astra, which it said likely marks the onset of AGI, led on computer use, coding and science and was its first model rated 'Critical' for cybersecurity under its Preparedness Framework; it rolled out first as a limited preview before reaching paid ChatGPT, the API, Azure and AWS.
- Google launches Gemini 3.8 Flash and Gemini 3.8 Flash CyberGoogle's third Flash release in six weeks improved software engineering, agentic tasks and multi-step reasoning over 3.7 Flash, alongside a Cyber variant gated to vetted defenders.
- Meta releases Muse Spark 1.3Muse Spark 1.3 improved agentic coding and long-horizon work while using ~20% fewer tool calls and ~25% fewer tokens than 1.2, shipping in Muse Code and the Meta Model API with closed weights.
- Anthropic releases Claude Fable 5.1 and Mythos 5.1The second generation of Anthropic's Mythos-class tier arrived three months after Fable 5, with cheaper cache reads cutting typical costs ~25% and safeguards producing 60% fewer cybersecurity false positives; Mythos 5.1 stays limited to trusted-access programs.
- DeepSeek V4-Pro leaves previewDeepSeek made V4-Pro generally available across its app, web and API with the agent-focused V4-Pro-0813 checkpoint, supporting 1M-token context and 384K-token outputs.
- Google releases Gemini 3.7 FlashArriving 23 days after 3.6 Flash, Gemini 3.7 Flash jumped from 49% to 65.3% on DeepSWE at half the launch price, while the promised Gemini 3.5 Pro remained delayed.
- SpaceXAI releases Grok 4.6Grok 4.6, built for long-running agents and coding with a 500K-token context, matched GPT-5.6 Sol on the Artificial Analysis Intelligence Index, one point behind Claude Fable 5.
- Meta releases Muse Spark 1.2 and Muse Code coding agentMeta Superintelligence Labs shipped the coding-focused Muse Spark 1.2 with Muse Code, a terminal coding agent that runs parallel sub-agents, and later committed to releasing Muse Spark 1.2's weights.
- Alibaba launches Qwen3.8-Max, a 2.4T-parameter flagshipQwen3.8-Max, a natively multimodal MoE with 2.4T total and 95B active parameters and a 1M-token context, became Alibaba's most capable model, with Alibaba claiming benchmark scores rivaling Anthropic.
- Anthropic releases Claude Opus 5Claude Opus 5 came close to Fable 5's frontier performance at half the price, topped the Artificial Analysis leaderboard at launch and added a user-selectable effort toggle.
- Google launches Gemini 3.6 FlashGemini 3.6 Flash improved coding, long-context and computer-use scores while using ~17% fewer output tokens than 3.5 Flash, launching with Gemini 3.5 Flash-Lite and a security-focused Gemini 3.5 Flash Cyber.
- OpenAI publicly releases GPT-5.6 Sol, Terra and LunaAfter limiting GPT-5.6 to trusted partners since June 26 at the US government's request, OpenAI opened its three-tier GPT-5.6 family to the public, led by the flagship Sol model for coding, science and cybersecurity.
- SpaceXAI releases Grok 4.5Grok 4.5, built on a 1.5-trillion-parameter base and pitched by Musk as 'Opus-class' but faster and cheaper, launched at $2/$6 per million tokens in the xAI API, Grok Build and Cursor.
- Anthropic releases Claude Sonnet 5Billed as the most agentic Sonnet, Claude Sonnet 5 performs close to Opus 4.8 at lower cost, with a 1M-token context, and became the default model for Free and Pro users.
- Anthropic releases Claude Fable 5, a public Mythos-class modelAnthropic launched Claude Fable 5, its first generally available Mythos-class model, with safeguards that fall back to Opus 4.8 in high-risk domains, alongside Claude Mythos 5 for vetted cybersecurity users.
- Microsoft unveils MAI-Thinking-1, its first flagship reasoning modelAt Build 2026 Mustafa Suleyman presented seven in-house MAI models led by MAI-Thinking-1, a ~1T-parameter MoE trained from scratch without OpenAI data that Microsoft says matches Opus 4.6 on SWE-Bench Pro.
- Anthropic releases Claude Opus 4.8Opus 4.8 improved on Opus 4.7 at the same price, led computer-use and browser-agent benchmarks, added dynamic workflows in Claude Code and cut Fast mode pricing by two-thirds; Anthropic said Mythos-class access was weeks away.
- Google I/O 2026: Gemini 3.5 Flash and Gemini OmniGoogle launched Gemini 3.5 Flash, which it said beats Gemini 3.1 Pro on coding and agentic benchmarks at Flash speed, introduced Gemini Omni models that output video, and shipped Antigravity 2.0.
- DeepSeek releases V4 preview with 1M-token contextDeepSeek open-sourced V4-Pro (1.6T parameters, 49B active) and V4-Flash (284B), using a new hybrid compressed attention that sharply cuts long-context compute and KV cache.Open source
- OpenAI releases GPT-5.5GPT-5.5 targeted long-horizon agentic, tool-using work and rolled out to paid ChatGPT and Codex users, with GPT-5.5 Pro and API access following a day later.
- Anthropic releases Claude Opus 4.7Opus 4.7 improved agentic coding, vision and instruction following at Opus 4.6 pricing, while Anthropic acknowledged it trails the unreleased Mythos.
- Meta Superintelligence Labs debuts Muse SparkMuse Spark, MSL's first model and Meta's first proprietary frontier model, is a natively multimodal reasoning model built in nine months under Alexandr Wang that powers the Meta AI app.
- Anthropic unveils Claude Mythos Preview and Project GlasswingAnthropic disclosed Claude Mythos Preview, a model a tier above Opus that autonomously found thousands of zero-day vulnerabilities, and withheld it from general release, instead giving access to defenders via Project Glasswing with partners including AWS, Apple, Google, Microsoft and Nvidia.
- Mistral releases Mistral Small 4Mistral Small 4, a 119B-parameter open MoE, unified the reasoning of Magistral, the vision of Pixtral and the agentic coding of Devstral in a single model.Open source
- OpenAI launches GPT-5.4 Thinking and ProOpenAI released GPT-5.4, billed as its most capable and efficient frontier model for professional work, with up to a 1M-token context in the API; mini and nano variants followed on March 17.
- Google releases Gemini 3.1 ProGemini 3.1 Pro more than doubled Gemini 3 Pro's ARC-AGI-2 score to a verified 77.1%, targeting complex reasoning and long-horizon agentic work.
- Anthropic releases Claude Sonnet 4.6Sonnet 4.6 brought near-Opus performance and a 1M-token context at Sonnet pricing, becoming the default model for Free and Pro users in Claude and Cowork.
- Alibaba releases Qwen3.5 open-weight multimodal modelQwen3.5, a 397B-parameter sparse MoE (17B active) native vision-language model supporting 201 languages, was released as open weights alongside a proprietary Qwen3.5-Plus, pitched at agentic AI.Open source
- Anthropic releases Claude Opus 4.6Claude Opus 4.6 led on agentic coding, computer use, search and finance, and introduced a 1M-token context window and agent teams in Claude Code.
- OpenAI releases GPT-5.3-CodexLaunched minutes after Opus 4.6, GPT-5.3-Codex combined the Codex and GPT-5 training stacks into a faster, steerable general-purpose coding agent model.
2025
35 models- Google launches Gemini 3 Flash as the default Gemini modelGemini 3 Flash brought near-Gemini 3 Pro performance at Flash latency and cost, becoming the default in the Gemini app and AI Mode in Search.
- OpenAI launches GPT-5.2 after internal 'code red'OpenAI released GPT-5.2 with gains in professional tasks, coding and long context, weeks after Sam Altman declared a 'code red' to respond to Google's Gemini 3.
- Mistral releases the Mistral 3 open-weight familyMistral launched Mistral Large 3, a 675B-parameter multimodal MoE (41B active), plus Ministral 3 dense models at 14B, 8B and 3B, all under Apache 2.0.Open source
- DeepSeek releases V3.2 and V3.2-SpecialeDeepSeek shipped MIT-licensed V3.2, its first model integrating thinking into tool use, and a high-compute V3.2-Speciale it claimed reached gold-level results at IMO and IOI 2025.
- Anthropic releases Claude Opus 4.5 at a 67% lower priceClaude Opus 4.5 reached 80.9% on SWE-bench Verified, ahead of GPT-5.1-Codex-Max and Gemini 3 Pro, while Opus pricing dropped to $5/$25 per million tokens.
- Google releases Nano Banana Pro (Gemini 3 Pro Image)Built on Gemini 3 Pro, Nano Banana Pro added legible multilingual text rendering, 4K output, search grounding and multi-image blending.
- Google launches Gemini 3 and the Antigravity coding platformGemini 3 Pro set records on Humanity's Last Exam and LMArena and shipped on day one in Search and the Gemini app, alongside the new agentic IDE Google Antigravity.
- xAI releases Grok 4.1Grok 4.1 focused on emotional intelligence and conversation quality, briefly taking the top spot on LMArena, with a 2M-token Grok 4.1 Fast variant.
- OpenAI releases GPT-5.1 Instant and ThinkingGPT-5.1 made ChatGPT warmer and better at following instructions, with adaptive reasoning and new personality presets, rolling out to all users.
- Anthropic releases Claude Haiku 4.5Claude Haiku 4.5 matched Claude Sonnet 4 on coding and computer use at one-third the cost and more than twice the speed.
- Anthropic releases Claude Sonnet 4.5Claude Sonnet 4.5 took the lead on SWE-bench Verified (77.2%) and computer use, sustained 30-hour autonomous coding runs, and shipped with Claude Code 2.0 and the Claude Agent SDK.
- Alibaba launches trillion-parameter Qwen3-MaxAt its Apsara Conference, Alibaba Cloud released Qwen3-Max, its largest model with over 1 trillion parameters trained on 36T tokens, targeting reasoning, coding and agents.
- Microsoft AI unveils its first in-house models, MAI-1-preview and MAI-Voice-1Microsoft AI introduced MAI-Voice-1 speech generation and began public testing of MAI-1-preview, its first end-to-end trained foundation model, signaling less reliance on OpenAI.
- Google launches Gemini 2.5 Flash Image, aka 'Nano Banana'Google revealed that the viral 'nano-banana' model topping image-editing leaderboards was Gemini 2.5 Flash Image, offering consistent characters, multi-image blending and natural-language edits.
- DeepSeek releases V3.1 hybrid thinking modelDeepSeek-V3.1 merged V3 and R1 capabilities into one 671B MoE model with switchable thinking and non-thinking modes and much stronger agent and tool-use performance.
- OpenAI launches GPT-5GPT-5 unified OpenAI's fast and reasoning models behind a real-time router, became the default model for all ChatGPT users including free tier, and set new marks in coding, math and health.
- OpenAI releases gpt-oss, its first open-weight models since GPT-2OpenAI released gpt-oss-120b and gpt-oss-20b under Apache 2.0; the larger model approaches o4-mini on reasoning benchmarks and runs on a single 80GB GPU.Open source
- Anthropic releases Claude Opus 4.1Claude Opus 4.1 upgraded Opus 4 on agentic tasks and real-world coding, reaching 74.5% on SWE-bench Verified at unchanged pricing.
- Alibaba open-sources Qwen3-CoderAlibaba released Qwen3-Coder-480B-A35B, an Apache 2.0 mixture-of-experts model for agentic coding with a 256K context, plus the Qwen Code command-line tool.Open source
- xAI launches Grok 4 and multi-agent Grok 4 HeavyxAI released Grok 4 and Grok 4 Heavy, a parallel multi-agent version, claiming top scores on Humanity's Last Exam and ARC-AGI-2, alongside a $300/month SuperGrok Heavy tier.
- Gemini 2.5 Pro and Flash reach general availability; Flash-Lite previewedGoogle made Gemini 2.5 Pro and 2.5 Flash stable and generally available and introduced the low-cost, low-latency Gemini 2.5 Flash-Lite in preview.
- Mistral releases Magistral, its first reasoning modelsMistral AI introduced Magistral, a reasoning model family with an open-weight 24B Magistral Small (Apache 2.0) and an enterprise Magistral Medium.
- OpenAI launches o3-pro and cuts o3 prices by 80%OpenAI released o3-pro, a higher-compute version of its o3 reasoning model, and cut o3 API pricing by 80% to $2/$8 per million tokens.
- Anthropic releases Claude 4 Opus and Claude 4 SonnetClaude 4 Opus and Sonnet launched with breakthrough coding and agentic capabilities, setting new state-of-the-art on SWE-bench.
- Anthropic releases Claude Sonnet 4 and Claude Haiku 4Alongside Claude Opus 4, Anthropic launches Sonnet 4 and Haiku 4 completing the Claude 4 model family.
- Google releases Gemini 2.5 FlashGemini 2.5 Flash released with built-in thinking capabilities, optimized for speed and cost efficiency.
- OpenAI releases o3 and o4-mini reasoning modelsOpenAI launches o3 (full release) and o4-mini with tool use capabilities including web browsing, code execution, and file analysis.
- Meta releases Llama 4 Scout and MaverickLlama 4 family released with MoE architecture: Scout (17B active/109B total) and Maverick (17B active/400B total) models.Open source
- Google releases Gemini 2.5 Pro with thinkingGemini 2.5 Pro released as Google's most capable model with built-in reasoning capabilities and thinking mode.
- OpenAI launches GPT-4o native image generationGPT-4o gains native image generation capability integrated directly into the model, producing highly realistic and text-accurate images.
- OpenAI releases GPT-4.5 research previewGPT-4.5 released as OpenAI's largest non-reasoning model, with improved EQ and reduced hallucinations.
- Anthropic releases Claude 3.7 SonnetClaude 3.7 Sonnet launched as a hybrid model with optional extended thinking, combining fast responses with deep reasoning.
- xAI launches Grok 3Grok 3 released as xAI's most capable model, trained on the Colossus supercomputer with 200k H100 GPUs.
- DeepSeek releases R1 reasoning modelDeepSeek R1 open-weight reasoning model released, matching o1 performance. Triggers global AI market shock and debate about AI cost assumptions.Open source
- OpenAI releases o3-minio3-mini launches as a cost-efficient reasoning model with adjustable reasoning effort levels.
2024
33 models- DeepSeek releases V3 open modelDeepSeek V3 (671B MoE) released open-weight, trained for reportedly ~$5.5M, competitive with top proprietary models.Open source
- OpenAI announces o3 reasoning modelOpenAI previews o3 and o3-mini, next-generation reasoning models achieving major improvements on ARC-AGI and math benchmarks.
- Google releases Veo 2 video and Imagen 3 image modelsGoogle launches Veo 2 for video generation and upgraded Imagen 3 for high-quality image creation.
- xAI open-sources Grok modelxAI releases Grok base model weights, making it one of the few frontier models available as open weights.Open source
- Google launches Gemini 2.0 FlashGemini 2.0 Flash released with native tool use, multimodal output, and improved performance over 1.5 Pro.
- OpenAI releases full o1 modelFull o1 model released to ChatGPT Plus users, with improved reasoning over the preview version.
- Alibaba releases Qwen 2.5 Coder and Qwen 2.5 72B instruct updatesQwen 2.5 Coder models released with state-of-the-art open-source code generation capabilities.Open source
- Anthropic releases upgraded Claude 3.5 Sonnet and Claude 3.5 HaikuUpdated Claude 3.5 Sonnet with major improvements in coding, plus new Claude 3.5 Haiku. Computer use feature launched in public beta.
- Meta releases Llama 3.2 with vision and lightweight modelsLlama 3.2 includes 11B and 90B multimodal vision models plus lightweight 1B and 3B text models for edge devices.Open source
- Alibaba releases Qwen 2.5 model familyQwen 2.5 released across multiple sizes (0.5B to 72B) with strong coding and math capabilities.Open source
- Mistral releases Mistral Small v2Updated Mistral Small model optimized for low-latency workloads with improved instruction following.
- Mistral releases Pixtral 12B multimodal modelPixtral 12B is Mistral's first multimodal model, handling both text and images natively.
- OpenAI releases o1 reasoning modelOpenAI introduces o1-preview and o1-mini, models trained with reinforcement learning to reason through chain-of-thought before answering.
- xAI releases Grok-2 and Grok-2 MiniGrok-2 launches with improved reasoning, competitive with frontier models on benchmarks.
- Mistral releases Mistral Large 2Mistral Large 2 (123B parameters) released with strong coding and multilingual capabilities.
- Meta releases Llama 3.1 with 405B modelLlama 3.1 family released including the 405B parameter model, the largest open-weight model at the time.Open source
- OpenAI releases GPT-4o miniSmall, affordable model replacing GPT-3.5 Turbo with significantly better performance.
- Mistral releases Mistral NeMo 12B with NVIDIA12B parameter model jointly developed with NVIDIA, Apache 2.0 licensed with 128k context.Open source
- Google releases Gemini 1.5 Flash & Pro GAGemini 1.5 models with 2M context and optimized latency/cost reach general availability.
- Google releases Gemma 2 open modelsGemma 2 models (9B, 27B) released with improved performance and efficiency.Open source
- Anthropic launches Claude 3.5 SonnetClaude 3.5 Sonnet released, outperforming Claude 3 Opus at faster speed and lower cost.
- DeepSeek releases Coder V2DeepSeek Coder V2 open model supports 338 programming languages and 128K context, competitive with GPT-4 Turbo on coding.Open source
- DeepSeek releases V2 MoEDeepSeek V2 mixture-of-experts model open-weight release.Open source
- Alibaba releases Qwen 2 model familyQwen 2 released with models from 0.5B to 72B, strong multilingual support and competitive performance.Open source
- OpenAI launches GPT-4oGPT-4o multimodal model with real-time voice and vision unveiled.
- Microsoft releases Phi-3 small language modelsPhi-3 Mini, Small, and Medium models released, achieving strong performance at small scale.
- Meta releases Llama 3 (8B & 70B)Llama 3 family open-weight models announced with improved benchmarks.Open source
- xAI announces Grok-1.5Grok-1.5 with long context released for xAI chatbot.
- Anthropic launches Claude 3 modelsClaude 3 Opus, Sonnet, Haiku released with strong reasoning and vision.
- Mistral releases Mistral Large and Le ChatMistral Large flagship model launched alongside Le Chat conversational assistant.
- Google releases Gemma open modelsGemma 2B/7B open models released with commercial-friendly terms.Open source
- Google announces Gemini 1.5 with 1M contextGemini 1.5 Pro announced with breakthrough 1 million token context window using mixture-of-experts architecture.
- Alibaba releases Qwen 1.5 modelsQwen 1.5 family (0.5B to 72B) released with multilingual support and improved performance.Open source
2023
11 models- Mistral releases Mixtral 8x7BMoE open-weight model with strong benchmarks, Apache 2.0 license.Open source
- Google launches Gemini 1.0Gemini Ultra/Pro/Nano announced as multimodal models for Bard/Pixel.
- Meta open-sources SeamlessExpressive/StreamingReal-time speech translation models released with code and weights.Open source
- OpenAI releases Whisper v3 and GPT-4 TurboAt DevDay, OpenAI launches GPT-4 Turbo with 128k context, Whisper v3, and custom GPTs.
- Mistral AI releases Mistral 7BFirst Mistral 7B dense model open-weight, strong performance for size.Open source
- OpenAI unveils DALL-E 3New image model with integrated ChatGPT prompting and better fidelity.
- Meta releases SeamlessM4TMultimodal translation model for speech and text, open for research.Open source
- Meta and Microsoft launch Llama 2Llama 2 open models (7B-70B) released for research and commercial use.Open source
- Anthropic releases Claude 2Claude 2 debuts with improved reasoning and 100k context.
- Google announces PaLM 2PaLM 2 unveiled at Google I/O for Bard and Workspace features.
- OpenAI releases GPT-4Multimodal GPT-4 model announced with API and ChatGPT Plus integration.