Models & Releases
New models, versions, and capabilities shipping across the frontier and open-model labs.
New models, versions, and capabilities shipping across the frontier and open-model labs.
Google has moved its next-generation text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, out of preview and into general availability on the Gemini API, alongside a new…
Google began rolling out "Call for Me" on September 24, 2026, a feature that lets Gemini place a phone call on a user's behalf and carry the conversation through to completion rather than just…
Anthropic has moved cache diagnostics on the Claude API out of beta, dropping the special header developers previously needed to turn the feature on. What's new The change is a straightforward…
ElevenLabs made Scribe v2 Medical generally available on September 22, 2026, a speech-to-text model fine-tuned specifically for clinical audio. It cuts word-error rate on medical dictation by roughly…
Google DeepMind published a technical update on September 23, 2026 describing a new server-side memory layer for Private AI Compute, its confidential-computing platform for personal AI assistants.…
OpenAI announced an upgraded prompt caching system for GPT-6, aimed at the multi-turn, long-running agent workloads the model family is built to handle. The update bundles a bigger discount, a wider…
OpenAI shipped two new models today, GPT-6 Sol and GPT-6 Luna, extending the GPT-6 family down into cheaper, more efficient tiers while cutting API prices in half compared with their GPT-5.6…
Anthropic released Claude Opus 5.5 on September 22, 2026, the first model in a new Claude 5.5 family that the company says matches its own Claude Fable 5.1 on most work while costing 40% less to run…
xAI shipped Grok 4.7 on September 21, 2026, calling it "our most capable model for coding and knowledge work" and pitching it as roughly twice as fast at half the price of comparable frontier models,…
Alibaba's Qwen team has released Qwen-Image-2.1, an open-weight image generation and editing model that folds text-to-image creation, reference-image conditioning, and localized editing into a single…
Anthropic has extended its Compliance API so enterprise administrators can retrieve transcripts of employee sessions with Claude in Chrome, the company's browser-based Claude agent, closing a…
Diogo Almeida, the former OpenAI researcher who helped build ChatGPT and co-invent reinforcement learning from human feedback (RLHF), released a new kind of AI model this week through his startup…
Google's Gemini API changelog recorded a new build of its Antigravity coding agent on September 17, 2026: antigravity-preview-09-2026, which replaces and deprecates the prior…
xAI has released Grok Voice Transcribe 2.0, a new speech-to-text model the company says is twice as accurate as its predecessor at unchanged pricing. The model is available now through xAI's API for…
AWS announced Amazon SageMaker HyperPod Inference Gateway, a Kubernetes-native routing layer for LLM inference clusters that the company says cuts first-token latency by as much as 82% without…
Alibaba's Qwen team launched Qwen3.8-Omni-Flash on September 14, 2026, a native omnimodal model the company says pushes omnimodal AI beyond "understanding" content toward planning tasks, calling…
Google Labs has expanded CC, its personal AI agent, from a single-user tool into a shared agent that families and households can use together. What's new CC now supports up to six people…
Suno launched v6 on September 9, 2026, a new generation of its AI music models built with input from major record labels, and retired every earlier Suno model in the process. What's new Suno…
xAI announced on September 16, 2026 that Grok Build, its coding agent, now has persistent memory: it automatically records project conventions, decisions, and facts as it works, then reads those…
Google has released two new real-time dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, expanding its live audio-and-vision model line for developers, enterprises, and consumer…
Anthropic has added on-demand context compaction to the Claude Messages API, letting developers shrink a long conversation into a signed summary block without losing the ability to keep recent turns…
Perplexity's Portable Computer — a local, on-device version of its Perplexity Computer agent — is now available on Windows PCs equipped with NVIDIA RTX GPUs, extending a feature that previously ran…
DeepSeek released V4.1-Flash on September 10, 2026, a new smaller model in its V4 line that the company says outperforms its own larger V4-Pro flagship on speed, cost, and total runtime. What's new…
Cursor released Projects on September 10, 2026, a new mode built around a single coordinator agent that plans and delegates work to as many subagents as a job requires, rather than writing code…
Anthropic has launched smart reports for Claude Enterprise, a beta analytics feature that turns raw usage data into a written report on how a team is actually using Claude — not just usage counts.…
Together AI has rolled out a significant expansion of its fine-tuning platform, adding support for a wider set of open-weight models alongside new tools for monitoring and controlling training jobs.…
Anthropic extended a beta feature that lets developers change a Claude model's reasoning effort mid-conversation to Google Cloud's Vertex AI platform on September 3, 2026. What's new Per the Claude…
OpenAI added a new account-security control to its API on September 10, 2026: the ability to force project API keys to expire automatically. What's new Per OpenAI's own changelog: "You can now set…
Anthropic has added a third permission policy, auto, to Claude Managed Agents, letting the server itself decide — call by call — whether an agent's tool use runs, gets denied, or pauses for human…
Google has released the Gemini app for Windows, bringing its AI assistant to the desktop as a standalone app rather than a browser tab or system tray widget. The app is available globally starting…
OpenAI made GPT-Live-1 generally available in the API on September 10, 2026, a dedicated voice model that listens and speaks at the same time and hands reasoning and tool use off to a backend model,…
OpenAI released the Agents API in public beta on September 10, 2026, giving developers hosted infrastructure to build and run production agents without assembling the harness themselves. The API…
OpenAI has introduced a Data agent inside ChatGPT Work that lets employees query company data, get investigations into what changed, and generate interactive, shareable dashboards through…
DeepSeek released DeepSeek-V4.1-Flash on September 10, replacing both DeepSeek-V4-Flash and the experimental vision-enabled variant it shipped in August with a single general-availability multimodal…
OpenAI has shipped three new Responses API controls aimed at long-running agentic work with GPT-6 Astra: async tool calling, mid-turn steering over WebSockets, and mid-conversation reasoning-effort…
OpenAI has added Prompt Cache Diagnostics to the Responses API, giving developers a way to see exactly why a request failed to reuse cached tokens instead of guessing. What's new The feature works by…
NVIDIA has released Nemotron-3-Labs-Ultra-Math-RL, the flagship 550-billion-parameter member of its Nemotron-3 family, specialized for mathematical reasoning. The model card, published on Hugging…
OpenAI shipped GPT Image 2.5 on September 8, 2026, replacing GPT Image 2 as the default image-generation model across ChatGPT and the API with two new variants: GPT Image 2.5 Flare and GPT Image 2.5…
Google has moved Gemini 3.8 Flash out of preview and into general availability, positioning it as the company's most capable Flash-tier model to date. The release is documented in the Gemini API…
Cursor has launched Self-Hosted Machines, a feature that lets its cloud coding agents execute on infrastructure an organization controls, while Cursor's cloud continues to handle inference and…
Microsoft AI shipped MAI-Transcribe-2 on September 3, 2026, a new speech-recognition model the company says is the fastest, most accurate, and cheapest transcription model on the market, priced at…
Google has put Lyria 3.5, its music generation model, into public preview on the Gemini API, extending a model that debuted in Google's Flow Music tool in July to developers building it directly into…
DeepSeek shipped the general-availability release of DeepSeek-V4-Pro on August 13, 2026, moving its flagship model out of preview with a significant upgrade to agent capabilities and native support…
Anthropic shipped version 1.30.0 of its ant command-line tool on September 3, 2026, adding a new ant apply command that manages Claude Platform resources — agents, environments, skills, memory…
OpenAI has begun shipping GPT-6 Astra, the model it describes as "our most intelligent model yet, with state-of-the-art performance in computer use, browsing, software engineering, science, and…
Google DeepMind and Google Research introduced WeatherNext 3 on September 3, 2026, a global weather forecasting model the companies say is now the most accurate available according to independent…
Meta rolled out Muse Spark 1.3 on September 2, 2026, the latest version of its Muse Spark model family, now live in Muse Code and Meta's own Model API. The release targets long-running agentic and…
Google shipped Gemini 3.8 Flash to general availability on September 2, 2026, alongside a restricted cybersecurity-focused sibling, Gemini 3.8 Flash Cyber, aimed at vulnerability discovery and…
Google announced agentic video understanding for Gemini on September 1, 2026, a new way for the model to analyze video by dynamically scanning segments instead of sampling frames at a fixed rate.…
Anthropic released Claude Fable 5.1 on September 1, 2026, the successor to Claude Fable 5 for long-running agentic coding, knowledge work, and research. Alongside it, Anthropic shipped an identical…
xAI has shipped a set of updates to the Grok Imagine image API, expanding the controls developers have over quality, source-image inputs, and output framing. What's new According to xAI's release…
Mistral has moved OCR 4.1, its document-intelligence model, out of preview and into General Availability, according to the company's changelog. What's new Mistral's changelog states plainly: "OCR 4.1…
Cohere has released Parse, a vision-language model purpose-built to turn complex enterprise documents into structured, machine-readable text. The company calls it "a high-throughput vision parsing…
Google has moved its fast, conversational video generation and editing model out of preview. On August 27, 2026, the Gemini API changelog announced the general-availability release of…
Google added a feature called Expert Intelligence to Gemini Notebook on August 27, 2026, letting users pull select ebooks they've bought from Google Play Books directly into a notebook so the AI's…
DeepSeek has released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal version of its V4 Flash model that adds image understanding on top of the model's existing text capabilities. What's new…
Anthropic rolled out two new API key types for the Claude Console on August 27, 2026: personal keys and service account keys, giving organizations a way to tie API usage to a specific person or…
ElevenLabs has moved its second-generation dubbing engine out of its consumer app and into the API, giving developers programmatic control over translating audio and video while preserving how each…
Google's Gemini Omni 1.1 Flash is now production-ready for developers, adding scene extension, first/last-frame control, and 4K upscaling to its generative video model — plus a cheaper 360p draft…
OpenAI has introduced Ultrafast mode, a new API service tier for its GPT-5.6 Sol model that runs dramatically faster than standard processing, according to the OpenAI API changelog. What's new The…
Google has taken its next-generation speech-to-text technology out of preview. In an August 26, 2026 update to the Gemini API changelog, Google announced that Gemini 3.5 Transcribe and Gemini 3.5…
Z.ai has confirmed that GLM-5.3-Flash, released August 26, is the model that spent the past week circulating under the code name "Ox Alpha" — an anonymous release that briefly became the top model on…
Google added a system-wide dictation mode to the Gemini app for macOS that transcribes spoken words into any desktop window while automatically stripping filler words and fixing mid-sentence…
AWS has added native Ray support to SageMaker HyperPod, letting data scientists spin up and manage distributed Ray clusters from the SageMaker console without hand-writing Kubernetes manifests.…
Anthropic has moved Claude Mythos 5 — its most cybersecurity-capable and most tightly restricted model — into Claude Security, the product that scans enterprise codebases for vulnerabilities and…
Anthropic has launched a new browser use tool for the Claude API, giving developers a way to have Claude drive a browser inside their own application rather than controlling an entire desktop. What's…
Mistral AI introduced Agentic Search on August 20, 2026, a retrieval system that lets its models iteratively search, inspect, and verify information across an organization's documents instead of…
OpenAI has released a Prompt Caching dashboard on the API platform, giving developers a direct view into cache hit rates, cache reads per write, and the token breakdown between cache-read,…
OpenAI now lets API customers pin an individual request to a specific processing region, using a prefixed domain on an account with Global-geography data settings — giving developers request-level…
Black Forest Labs, the German AI lab behind the FLUX image and video model family, launched FLUX Video Upscale on August 20, 2026 — a standalone API endpoint that upscales video to 1080p, 2K, or…
DeepSeek released a new experimental vision-language model on its API platform on August 21, 2026, extending the V4-Flash line to multimodal image understanding for the first time. What's new The…
OpenAI has added native transparent-background output to gpt-image-2, its flagship image generation model, according to an entry posted to the OpenAI API changelog on August 20, 2026. The capability…
Amazon Web Services has added cross-Region inference for OpenAI's GPT-5.6 model family on Amazon Bedrock, letting requests route across multiple AWS regions based on available capacity rather than…
Anthropic has simultaneously graduated three previously-beta pieces of the Claude Platform to general availability: the Files API, Admin API user-management endpoints for Claude Enterprise…
Vercel has built an autonomous agent pipeline, internally called ai-sdk-factory, that now authors more than a quarter of the pull requests merged each week into the AI SDK, one of the most widely…
Anthropic has renamed the Claude Console's prompt-testing tool from Workbench to Playground, pairing the rebrand with expanded coverage of the Messages API and new visibility into the raw requests…
Amazon Web Services has taken Bedrock AgentCore payments out of preview, announcing general availability on August 18, 2026 for the service that lets AI agents autonomously pay for APIs, MCP tool…
DeepSeek has taken its V4-Pro-0813 model out of preview and into general availability across the app, web, and API, according to the company's official API changelog. The release keeps the same…
Z.ai has released GLM-5.3, a new flagship version of its GLM model line focused on complex software engineering and long-horizon agent tasks, according to the company's own documentation. The model…
Mistral has put OCR 4.1, the latest version of the OCR engine behind its Document AI stack, into public preview, adding native paragraph-level bounding box extraction and confidence scoring at the…
Cursor has shipped a feature it calls "Builds" for Cloud Agents, giving each coding agent a pre-configured, warm copy of a project's environment instead of setting one up from scratch every session —…
Google is rolling out a setting that lets Gemini users disable the visible watermark that has stamped every AI-generated image, video, and song, while keeping the invisible tracking layer intact…
Google has released Gemini 3.7 Flash, the latest entry in its low-latency Flash line, positioning it as the company's strongest model yet for software development and autonomous agent work. What's…
OpenAI opened a limited preview of Ultrafast mode on August 13, 2026, a new API service tier that runs GPT-5.6 Sol at up to 14 times the speed of standard processing by routing inference through…
DeepSeek has quietly moved its flagship language model out of preview status, updating its official API documentation to show the deepseek-v4-pro endpoint now running build DeepSeek-V4-Pro-0813. The…
OpenAI has taken its ChatGPT advertising program international, confirming on August 11, 2026 that ChatGPT Ads is now live in the United Kingdom, Mexico, Brazil, Japan, and South Korea. The move…
Google DeepMind has introduced a new sign-language AI model that translates sign language directly into text, and shipped it as a real product feature on the Pixel 11 rather than a research demo.…
Google unveiled the Pixel 11, Pixel 11 Pro, and Pixel 11 Pro XL at its Made by Google 2026 event on August 12, pairing modest hardware upgrades with a deeper push of Gemini Intelligence into everyday…
xAI shipped Grok 4.6 on August 12, 2026, the latest entry in its frontier model line aimed at coding, agentic tasks, and knowledge work. The model is live now on the xAI API with a larger context…
Google is rolling out a set of AI-generated reporting and insight features across Google Ads and Google Analytics, aimed at surfacing performance trends without marketers having to build reports…
Anthropic will make "auto mode" the default behavior for Claude Code on Pro, Max, and Team plans starting August 14, 2026, ending the practice of prompting developers to approve nearly every action a…
OpenAI expanded its Fast mode API tier on August 5, 2026 to cover long-context requests on the GPT-5.6 family, letting developers get faster responses on prompts that previously had to run at…
Anthropic shipped four updates to Claude Managed Agents on August 7, giving developers hard spend caps, a way to have one model consult another mid-task, control over where inference runs, and…
OpenAI began rolling out an August refresh of GPT-5.6 Sol and GPT-5.6 Luna in ChatGPT on August 6, 2026, replacing GPT-5.5 Instant as the default models across every ChatGPT tier and adding a new…
DeepSeek has moved the V4-Flash API out of preview and into public beta, pairing the shift with a fresh round of post-training that produces sizable jumps on agentic coding and terminal benchmarks.…
Anthropic has opened a beta of Inference Hooks for Claude Enterprise organizations, a new governance layer that lets companies route every prompt through their own security infrastructure before…
Cursor has launched three Google Workspace plugins — for Gmail, Google Drive, and Google Calendar — that let its coding agents read, write, and act on a developer's email, files, and schedule without…
Alibaba unveiled Qwen3.8-Max on August 3, 2026, the largest model in its Qwen family to date, with the company saying its benchmark scores are competitive with Anthropic's frontier models. The model…
OpenAI is retiring Priority Processing in favor of a new API tier called Fast mode, announced in the same July 30 update that cut prices on GPT-5.6 Luna and Terra. For GPT-5.6 Sol, the company's top…
Midjourney released its V8.2 image model on July 24, 2026, focused on aesthetic quality, reducing low-quality outputs, and sharpening how the tool learns individual users' taste. What's new…
xAI has expanded Grok Imagine Video 1.5 with reference-based video generation, letting users lock a specific face, product, or location into a generated clip, alongside a jump to native 1080p output…
Google has added direct Chrome integration to Gemini Spark, its autonomous personal AI agent, letting the agent use a person's logged-in accounts and saved passwords to complete web-based errands on…
DeepSeek moved DeepSeek-V4-Flash out of preview and into an official public beta on July 31, re-post-training the same architecture to deliver what the company describes as significantly enhanced…
NVIDIA expanded its Agent Toolkit on July 26, 2026 to include NVIDIA PhysicsNeMo and CUDA-X libraries as agent-callable tools, letting AI agents invoke physics simulations and accelerated solvers…
Google DeepMind introduced Gemini Robotics 2 on July 30, 2026, describing it as "the intelligence layer powering the next generation of truly adaptable robots." The release pairs a new…
xAI has released Grok Voice Think Fast 2.0, the successor to its real-time voice model, and will automatically migrate all traffic pointed at the grok-voice-latest alias to the new version starting…
Google DeepMind has released Lyria 3.5, the latest version of its music-generation model, rolling it out today inside Google Flow Music. The update focuses on four areas of the music-creation…
Google has added a new voice-driven natural-language mode to the Gemini app for macOS, letting users dictate, summarize, and rewrite text in any desktop application by holding a single key. What's…
AWS announced on July 28, 2026 that AgentCore Gateway now supports the 2026-07-28 revision of the Model Context Protocol (MCP), which the post describes as the protocol's biggest change since it…
xAI introduced Build Mode on July 28, 2026, a new feature inside Grok that turns a plain-language description into a working, publishable app, game, website, or dashboard without the user writing or…
OpenAI added two new speech-to-text models to its API on July 28, 2026: GPT Transcribe for file-based transcription and GPT Live Transcribe for low-latency streaming transcription, rounding out a…
Google has moved Gemini 3.6 Flash and Gemini 3.5 Flash-Lite out of preview and into general availability on the Gemini API as of July 21, 2026, while simultaneously deprecating the classic sampling…
Cursor has launched Router, a feature that automatically picks which AI model handles each coding request instead of leaving that choice to the developer, aiming to cut costs without giving up…
Black Forest Labs introduced FLUX 3, a multimodal foundation model the company says jointly learns from images, video, and audio inside one architecture rather than stitching together separate…
Anthropic released Claude Opus 5 on July 24, 2026, the newest version of its mid-tier flagship reasoning model, available immediately at the same price as its predecessor. What's new According to…
xAI updated its Speech to Text API with an adjustable voice-activity-detection threshold, giving developers finer control over how the model distinguishes speech from silence or background noise.…
OpenAI has begun rolling out Health in ChatGPT to logged-in US users 18 and older on web and iOS, across the Free, Go, Plus, and Pro plans. The feature lets people securely connect Apple Health and…
Anthropic launched a Claude connector for its Economic Index on July 22, 2026, letting anyone query the organization's dataset on real-world AI usage directly through conversation instead of digging…
AWS has published details on Agentic Retrieval, a new capability for Amazon Bedrock Managed Knowledge Base that lets a foundation model decompose complex, multi-part questions into sub-queries,…
Anthropic shipped a batch of Claude Managed Agents API updates on July 22, giving developers finer control over agent cost, faster session startup, and webhook-driven visibility into agent…