Updated

AI models released in September 2026

42 model releases from 23 providers in September 2026, with announcement dates, official sources, and practical notes for builders.

Releases this month

Dates refer to the linked announcement or release milestone. Previews are labeled in the entries; historical availability and pricing may have changed.

Release details

  1. Cohere

    Embed 5 Pro and Embed 5 Fast

    Cohere released Embed 5, a family of multimodal embedding models in Pro and Fast tiers. Both support text and image inputs, more than 100 languages, and a 128K-token context; their shared embedding space lets teams index with Pro and query with either tier.

    For buildersTest both tiers on representative text, image-rich documents, and multilingual queries; compare retrieval quality, latency, and index consistency before choosing production roles.

  2. Google

    Gemini 4 Argon (limited access)

    Google announced Gemini 4 Argon, a model designed for long-horizon reasoning across software engineering, enterprise knowledge work, and cybersecurity defense. Its initial rollout is limited to trusted cyber defenders through the Fairwind Program while Google continues guardrail work before broader access.

    For buildersIf you have access, compare Argon with your current model on a representative long-horizon coding or document-work task, tracking completion quality and review effort rather than relying on provider benchmark claims.

  3. Perplexity

    pplx-embed-v2-context-9b-preview

    Perplexity introduced pplx-embed-v2-context-9b-preview, a contextual embedding model that encodes document chunks with document-wide context and returns one embedding per chunk. It is publicly available as a preview; the provider cautions that weights and interface may change and that preview embeddings should not be mixed with later versions.

    For buildersTest it on long documents where answer passages depend on distant context, comparing retrieval of both answer-bearing and supporting passages against your current chunk embeddings.

  4. OpenAI

    GPT-6.1 Sol

    OpenAI released GPT-6.1 Sol, an upgrade to GPT-6 Sol that it says nearly matches GPT-6 Astra on agentic coding, computer use, and professional work at one-fifth of Astra's standard token price. It is available today in ChatGPT Work and Codex and through the API as gpt-6-1-sol.

    For buildersUse GPT-6.1 Sol for everyday coding and agent workloads and escalate to GPT-6 Astra only when tasks fail, since Sol costs $2 input and $10 output per million tokens versus Astra's $10 and $50.

  5. NVIDIA

    NVIDIA Kumo Tabular

    NVIDIA released Kumo Tabular, an open foundation model for tabular classification and regression. It predicts labels for new rows from labeled examples and is available in Small, Medium, and Large sizes.

    For buildersEvaluate it on a held-out tabular task against a strong tree-based baseline, especially when labeled examples are limited.

  6. AssemblyAI

    Universal-3.6 Pro Realtime

    AssemblyAI released Universal-3.6 Pro Realtime, an upgrade to Universal-3.5 Pro for streaming speech recognition. It supports 32 languages with automatic detection and is designed to handle short responses, multiple speakers, accents, and noisy telephony audio at the previous version’s latency.

    For buildersTest it on a representative set of noisy calls with short confirmations, names, and code-switching, and measure transcription errors before connecting its output to automated actions.

  7. Voltropy

    Vast-10M (early access)

    Voltropy announced its first model family, Vast-10M, in Flash, Medium, and Pro variants, using its Voltropy Scalable Attention method to extend transformer context. The company is offering early access; general availability is not confirmed.

    For buildersIf admitted to early access, test recall on a large document or code corpus against a smaller-context baseline.

  8. Anthropic

    Claude Sonnet 5.5

    Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family, positioned as a faster and lower-cost complement to Opus 5.5 for everyday coding and document work. It is priced the same as Sonnet 5 at $2 per million input tokens and $10 per million output tokens and generates output more than 30% faster.

    For buildersTry Sonnet 5.5 as a faster default for well-scoped coding and document tasks, and measure cost per finished task on your own workloads to confirm the claimed token savings.

  9. Hcompany

    Holo4

    Hcompany released Holo4, a family of agentic models in two sizes: a 27B dense model and a 35B MoE (35B-A3B) model, plus an updated Holotron4 Nano. They work across GUIs, code, MCP, and APIs, and are available through the H Models API with open weights on Hugging Face.

    For buildersEvaluate Holo4 through the H Models API on a multi-step task that mixes screen interaction, code execution, and API tool calls, since it is built to switch between interfaces within one workflow.

  10. Kuaishou

    Kling 4.0 (preview)

    Kling AI announced Kling 4.0, its next-generation video generation model, with single clips up to 30 seconds, 4K and 1080p 10-bit HDR output, and up to 15 multimodal reference inputs. Kling 4.0 Flash opened limited early access on September 28, and the full release is planned for October.

    For buildersWhen Kling 4.0 reaches general availability, test its multi-keyframe storyboard control for long-form and commercial video work, and compare per-generation cost and clip length with ByteDance Seedance 2.5.

  11. AutoTrust AI

    JEV-27B

    AutoTrust AI introduced JEV-27B, an open-weights model that pairs calibrated, typed one-pass decisions with a Qwen3.8-27B text-generation and reasoning path. The project publishes the weights under Apache-2.0.

    For buildersEvaluate it for triage or routing by checking calibration on held-out labeled examples and the quality of the generation path.

  12. Meituan

    LongCat-2.5 Preview

    Meituan launched LongCat-2.5 Preview on its API platform, adding native image understanding to its long-context model. It keeps a 1.6T parameter MoE architecture with about 48B active parameters, a 1M token context window, and 128K token output, and targets long-horizon agent tasks across browsers, terminals, and desktop apps.

    For buildersTest LongCat-2.5 Preview through Meituan's API for long-context agent work such as reading and modifying a large codebase, since its endpoints are compatible with Anthropic and OpenAI style calls.

  13. Perceptron

    Perceptron Mk1.5

    Perceptron released Mk1.5, a model for controlling embodied agents including drones, robotic dogs, smart glasses, and phones. It adds native audio input, video object tracking with timestamped geometries, and egocentric video understanding, with a 32K token context window.

    For buildersTry it through the Perceptron Platform API at $0.15/M input tokens for tasks that require real-time video tracking or embodied agent control without platform-specific retraining.

  14. CLM-8B

    Stanford and NVIDIA researchers released CLM-8B, an open contrastive language model that selects the best matching action for a state instead of generating tokens, which makes it fast for bounded agent decisions. It is available under Apache 2.0 with open weights and targets computer use, tool calling, and verification tasks.

    For buildersUse CLM-8B as a fast action selector or verifier in agent pipelines where the model chooses between a fixed set of tools or candidate outputs, and cache action embeddings to cut latency on repeated decisions.

  15. Black Forest Labs

    FLUX 3 Action

    Black Forest Labs released FLUX 3 Action, an open-weight 7B world action model that predicts future actions from camera frames and text instructions. Fine-tuned on DROID it placed first on the RoboLab-120 benchmark at 42.92% success rate, beating models twice its size with 3.95x faster inference.

    For buildersUse FLUX 3 Action through the LeRobot integration to fine-tune a robotics policy for your own robot arm or test it in simulation through NVIDIA Isaac Sim before deploying on hardware.

  16. Google

    Gemini 3.8 Flash TTS and Flash-Lite TTS

    Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS, expressive text-to-speech models that create custom voices from prompts, replicate a voice from a 30-second sample, and control delivery line by line. Both are available today in Google AI Studio and the Gemini API, with Flash-Lite aimed at high-volume dubbing and localization.

    For buildersPrompt a custom voice in Google AI Studio and rehearse line-by-line delivery in the dual-speaker editor before wiring it into your app through the Gemini API.

  17. Meta

    Muse Realtime Avatar

    Meta introduced Muse Realtime Avatar, an embodiment model that turns Muse Realtime Voice into expressive interactive avatars with synchronized lip movement and gestures. It powers live video conversations in the Muse app, which Meta says are rolling out, and generated video is watermarked with Video Seal.

    For buildersRun a short live conversation with a reference photo or illustration and check whether the avatar keeps consistent appearance and lip sync across turns before building a branded character around it.

  18. NVIDIA

    Nemotron 3 Diarization

    NVIDIA released Nemotron 3 Diarization, a 100-million-parameter open-weight model that identifies who spoke when in real time for up to eight speakers, independent of speech recognition. NVIDIA reports it ranks first in VoiceArena's initial Diarization-Bench with a 14.72 percent diarization error rate.

    For buildersDownload the open weights and pair its speaker timestamps with a transcription model to build meeting notes or call summaries without sending audio to a cloud API.

  19. NVIDIA

    NV-Reason-CT

    NVIDIA released NV-Reason-CT, an open research vision language model that applies chain of thought reasoning to 3D CT scans. It generates structured reports and supports multi-step follow-up questions for chest and abdominal CT, and is positioned as a foundation model for researchers rather than a cleared clinical product.

    For buildersYou can post-train NV-Reason-CT on the open weights for CT-specific applications, but it is a research foundation model, not approved for clinical diagnostic use.

  20. OpenAI

    GPT-6 Sol

    OpenAI released GPT-6 Sol for coding and professional work, with API pricing of $2 per million input tokens and $10 per million output tokens. Access is rolling out in ChatGPT Work and Codex; it is not yet available in Chat.

    For buildersEvaluate a coding or research workflow on your own tasks, comparing quality and total cost with the model you currently use.

  21. Anthropic

    Claude Opus 5.5

    Anthropic released Claude Opus 5.5, the first model in its Claude 5.5 family, for long-running coding and knowledge work. Company pricing is $4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 per million.

    For buildersRun the same long coding or knowledge-work job on Opus 5.5 and Opus 5, then compare total token cost and whether the cheaper cache-read rate actually shows up in the bill.

  22. OpenAI

    GPT-6 Luna

    OpenAI released GPT-6 Luna alongside Sol as a lower-cost model, with API pricing of $0.10 per million input tokens and $0.50 per million output tokens. Free and Go users can access Luna in the desktop app.

    For buildersTest a high-volume workflow such as classifying support requests or extracting structured fields from incoming documents.

  23. Tencent

    Hy Image 3.5 Preview

    Tencent released Hy Image 3.5 Preview, an image generation model supporting text-to-image, image-to-image with up to 5 reference images, and multi-turn conversational editing at up to 2K resolution. It is available through Tencent Cloud TokenHub API at $0.024 per image.

    For buildersIntegrate Hy Image 3.5 Preview's messages-based API for iterative image editing workflows where users refine visuals through conversation rather than one-shot generation, useful for e-commerce and content production pipelines.

  24. Xiaomi

    MiMo-V2.6-Pro and MiMo-V2.6-Flash

    Xiaomi open-sourced the MiMo-V2.6-Pro and MiMo-V2.6-Flash multimodal LLMs, with Pro scoring 46 on the AI Composite Intelligence Index to become the top open-weight model globally. Both models use scaled reinforcement learning for recursive self-improvement and are available as open weights.

    For buildersTry MiMo-V2.6-Pro or MiMo-V2.6-Flash through their open-weight releases for cost-efficient multimodal tasks like image understanding, 3D perception, and computer use, especially if you want to avoid per-token API costs from closed-source providers.

  25. Qwen

    Qwen-Audio 3.1 ASR, TTS, and Realtime

    Alibaba updated its Qwen-Audio 3.1 speech suite at Apsara 2026, covering automatic speech recognition, text-to-speech, and real-time voice interaction models with enhanced voice AI capabilities.

    For buildersWire the Realtime model into a voice agent prototype and test turn-taking and interruption handling in a noisy room before scaling it.

  26. Qwen

    Qwen-Audio-3.1-TTS-Next

    Alibaba unveiled Qwen-Audio-3.1-TTS-Next at Apsara 2026, an audio generation model that creates cinematic soundscapes by blending dialogue and ambient sound from a text script. It targets audiobooks, film, TV, podcasts, and games.

    For buildersFeed it a short script with dialogue and scene cues and compare the mixed audio against separate speech and effects pipelines for your project.

  27. Qwen

    Qwen-Image 3.1 (announced)

    Alibaba announced Qwen-Image 3.1, an image generation model for design and e-commerce marketing with native transparent-background generation and image editing. It is scheduled to launch later this year.

    For buildersPlan a hero-image and product-photo workflow around transparent-background generation once it ships, and compare it against your current image model.

  28. Qwen

    Qwen3.8-LiveTranslate

    Alibaba debuted Qwen3.8-LiveTranslate at Apsara 2026, a simultaneous interpretation model for real-time translation designed for smoother, more responsive live translation. It launched alongside the Qwen-Audio 3.1 speech and image model updates.

    For buildersTest it against a live bilingual call or meeting feed and compare latency and fluency with a streaming translation API before committing.

  29. Qwen

    Qwen4 family (announced)

    Alibaba previewed the Qwen4 family at Apsara 2026, naming Qwen4-Max, Qwen4-Flash, Qwen4-Plus, and Qwen4-27B. The models are still in training with no release date, specs, or weights published yet, and the Qwen 4.5 and Qwen 5 series are planned at 5 to 10 trillion parameters.

    For buildersKeep building on Qwen3.8-Max for now and watch the 27B tier if you want a locally deployable multimodal option when Qwen4 ships.

  30. xAI

    Grok 4.7

    xAI released Grok 4.7, a coding and knowledge-work model listed at the same API price as Grok 4.6: $2 per million input tokens and $6 per million output tokens. It is available in Cursor, Grok Build, and the xAI API.

    For buildersRun a long coding or document task on Grok 4.7 in Cursor or the API at Grok 4.6 pricing and compare quality with the model you currently use.

  31. Qwen

    Qwen-Image-2.1

    Qwen released Qwen-Image-2.1, an open-source image model that combines text-to-image generation and image editing in a single pipeline with native transparent image support and multi-image editing.

    For buildersTry generating a product image with an alpha channel in a no-code design tool, then edit it locally using reference images in one model workflow.

  32. Qwen

    Qwen3.8-Omni-Flash

    Qwen released Qwen3.8-Omni-Flash, an omnimodal AI model with 1M-token context that processes text, image, audio, and video inputs with agentic task planning and tool use capabilities.

    For buildersTest the model on a multi-step video editing or meeting summarization workflow to evaluate its omnimodal planning and tool use reach.

  33. StepFun

    Step 5 Preview

    StepFun released Step 5 Preview, its flagship agentic model built on a sparse MoE architecture with 600B total parameters (27B active) and a 1M token context window. It targets software engineering, professional knowledge work, and finance, and is available via the StepFun API. Open weights are expected by October 15, 2026.

    For buildersEvaluate it for long-horizon coding and agentic workflows through the StepFun API. The 1M context window makes it suited for tasks that require sustained context across many turns.

  34. xAI

    Grok Voice Transcribe 2.0

    xAI released Grok Voice Transcribe 2.0, a speech-to-text model available through its Speech-to-Text API for batch and streaming transcription. Company-listed pricing is unchanged from version 1.0: $0.10 per hour of audio for batch and $0.20 per hour for streaming.

    For buildersSend a noisy support call or meeting recording through the Speech-to-Text API and compare speaker labels, timestamps, and formatted emails or phone numbers against your current transcriber.

  35. Google

    Gemini 3.8 Flash

    Google announced 3.8 Flash with improvements to coding, reasoning, and multi-step agent workflows, alongside a separate Cyber variant.

    For buildersEvaluate it on your app's real tasks: extracting data, routing support requests, or carrying out a sequence of tool calls.