AI DevelopmentPricing Tracker5 min readPublished October 1, 2026

A living ledger, dated from the vendor’s page, refreshed through October 31

AI Model Releases: October 2026 Tracker and Dated Ledger

Every AI model released in October 2026, dated from the vendor's own page: price per million tokens, context window, licence and where you can use it today.

DA
Digital Applied Team
Research and practical guidance
CoverageOctober 1–3, 2026

October 2026 opened without a frontier launch and with a cluster of small, specialised models instead. On October 1 Cloudflare and Amazon each released open-weight decision models, Microsoft AI released two text-to-speech models and a streaming transcription model, and Tavus previewed a full-duplex video model to invited testers. On October 2 Bilibili’s Index team published the preview of a 35B-total, 3B-active translation model. This page records each one on the vendor’s own date and is refreshed in place through the month.

Editorial note: Opened October 3 as an October 1, 2026 dispatch. Rows and listings dated October 2 postdate the dateline and were collected on October 3 with the rest; rows added after October 3 carry the date they were added. Facts are as of October 3 unless a row says otherwise.

Key takeaways
  1. 01
    Eight models, five vendors, no frontier LLMEvery release in the first three days is a specialist: decision, speech, video or translation. The frontier news is an availability change, not a launch.
  2. 02
    Open weights outnumber closedFour of the eight carry Apache 2.0 licences with public weights. Microsoft’s three and Tavus’s one are API or invite only.
  3. 03
    Three listings, no vendor pageThree models appeared on OpenRouter on October 1 and 2 with no release page we could locate. They are recorded as listings, not releases.
  4. 04
    Argon is still not publicGemini 4 Argon, announced September 30, is reaching trusted cyber defenders first. Its introductory price is published; its general date is not.
Releases
On a vendor page, Oct 1–3
8models

Counted by model, so Clef and Clef-flash are two.

Vendor-dated
Open
Apache 2.0 with public weights
4of 8

Clef, Clef-flash, Strands Decider 2B and Index-Translate-35B.

Hugging Face
Listings only
On OpenRouter without a release page
3models

Recorded in section 04 and excluded from the counts above.

Corroboration only

01 — The ledgerThe October ledger

One row per model, dated from the vendor’s announcement. Where a vendor released two sizes in one post they share a row. The primary column names the page each row was read from; the linked sources are collected in the method section.

AI model releases dated October 1–3, 2026, as of October 3. Dates are the vendor’s own; OpenRouter listing times are corroboration only and appear in section 04. Griffin-Lite’s vendor pages are named rather than linked.
DateVendor · modelWhat it isPrimary source
Oct 1Cloudflare · Clef and Clef-flashDecision models (typed answers with probabilities, no text), 27B and 9B, post-trained from Qwen bases; read text, JSON, images and videoCloudflare blog, October 1; Hugging Face cards modified October 1
Oct 1Amazon Strands Labs · Strands Decider 2B2B decision model for local CPU or GPU; weights, training data and scripts releasedStrands blog, October 1; Hugging Face card modified October 1
Oct 1Microsoft AI · MAI-Transcribe-2-StreamingStreaming speech-to-text, 60 languages, continuous language detectionMicrosoft AI news post, October 1
Oct 1Microsoft AI · MAI-Voice-2.1Text-to-speech, 23 languages and 26 locales, one voice across languagesMicrosoft AI news post, October 1
Oct 1Microsoft AI · MAI-Voice-2.1-FlashLow-latency variant of the above for high-volume use; vendor states 150 ms end to endMicrosoft AI news post, October 1
Oct 1Tavus · Griffin-LiteResearch preview of a full-duplex video-to-video conversation model; the full Griffin model is announced for laterTavus announcement page and Business Wire release, October 1 (not linked)
Oct 2Bilibili Index team · Index-Translate-35B-A3B-previewMixture-of-experts translation model, 35B total and 3B active parameters, 150 text languages; dense 2B and 9B siblings reached Hugging Face on September 28Hugging Face card created October 2; Index-Translate project site

The two decision-model rows are the month’s first story, and they have their own page: open decision models compared puts Clef, Clef-flash and Decider 2B beside TypeSafe’s Jev on licence, size, context and the vendor’s benchmark rows.

02 — The termsPrice, context and licence

Prices as the vendor lists them on October 3, 2026: per million input tokens for the text models, per million characters for Microsoft’s speech models, per hour of audio for transcription. Context in tokens where the vendor states it.
ModelPriceContextLicence and where it runs
Clef / Clef-flash$0.24 / $0.09 per million input tokens on Workers AI65,536Apache 2.0; Workers AI or self-hosted
Strands Decider 2BFree, localNot statedApache 2.0; GitHub and Hugging Face
MAI-Transcribe-2-Streaming$0.54 per hour of audio, introductory through end of 2026n/aClosed; Microsoft AI and Azure
MAI-Voice-2.1$22 per million charactersn/aClosed; Microsoft AI, Azure, OpenRouter
MAI-Voice-2.1-Flash$15 per million charactersn/aClosed; Microsoft AI, Azure, OpenRouter
Griffin-LiteNot publishedn/aClosed; select early testers only
Index-Translate-35B-A3B-previewFree weights262,144 in the shipped config; card examples serve 32,768Apache 2.0; Hugging Face and ModelScope

Microsoft’s post attaches two claims to the Flash voice model that are the vendor’s own and unverified here: that inference is 55% faster and roughly 60% cheaper than comparable models, and that it produces 45 seconds of audio at 150 milliseconds end to end. Our September ranking of text-to-speech models is where the two voice models will be scored once listener data exists.

03 — ContextAvailability changes, not releases

The frontier model in the news this week is not in the ledger, because it was announced in September and is not yet generally available. Google announced Gemini 4 Argon on September 30 as rolling out to a set of trusted cyber defenders through its Fairwind programme, with an introductory price of $2 per million input tokens and $10 per million output tokens and cached input at 95% off. The post gives no date for developers, enterprises or consumers. Our Argon page tracks that date.

Three September releases sit just outside this ledger and are not counted here: Together’s Tev1 4B experimental, published on GitHub and Hugging Face on September 23 and listed on OpenRouter on September 30; Inception’s Mercury Decide, listed on OpenRouter on September 30 with no vendor page located; and Bilibili’s dense Index-Translate 2B and 9B, on Hugging Face since September 28, which the October 2 preview extends. The two decision models are covered in decision models that return probabilities, not text. The month before this one is in the September tracker.

04 — CorroborationListings without a vendor page

Recorded, not counted

OpenRouter’s model list added three entries in the window with no vendor release page located by October 3: Unbiased’s Pareto 26.10 Preview on October 1 (1.05M context, $0.80 input and $3.20 output per million tokens), Apodex 1.1 Mini on October 1 (262K, free tier), and inclusionAI’s Ling 3.1 Flash on October 2 (262K, free). A listing time is a listing, not a launch. Each moves into the ledger when a vendor page with a date appears.

05 — AnalysisWhat the first days show

The pattern is the inverse of September’s opening, which brought three frontier models in two days. This month began with models that do one thing: decide, speak, transcribe, translate, or hold a face-to-face conversation. Two of the five vendors are infrastructure companies, Cloudflare and Amazon, releasing weights rather than selling tokens. That is consistent with where decision models sit in an agent: close to the tool call, where a hosted round trip is the cost.

For a buyer, the practical reading is that none of the October rows replaces a model already in production. They add a new slot: a cheap check before an action, a voice that keeps its identity across languages, a translator with open weights. The question to ask of each is not “is it better than the frontier” but “what does it let me stop sending to the frontier”. Our AI transformation work starts from that question.

06 — MethodMethod and as-of date

Methodology

A census of vendor announcements. No model on this page was run by Digital Applied.

What was collected
Every AI model with a vendor-dated release page between October 1 and October 3, 2026, with its price, context window, licence and hosting as stated by the vendor. One row per model, except where a vendor released two sizes in one announcement.
Sources
Cloudflare’s Clef announcement and Workers AI model pages; Amazon’s Strands Decider post on the Strands Agents site, dated October 1; Microsoft AI’s transcription and voice post on its news site, dated October 1; the Index-Translate-35B-A3B-preview card; Hugging Face API records for licence tags, base-model tags and creation dates; Tavus’s announcement page and its Business Wire release; the OpenRouter models API for listing times.
As-of date
October 3, 2026. The ledger was opened on that date and covers October 1–3. Later rows carry the date they were added.
Dating rule
A row’s date is the vendor announcement date. A Hugging Face card created before the announcement does not move the date earlier; an OpenRouter listing never sets a date on its own.
Exclusions
Availability changes to models announced earlier, such as Gemini 4 Argon; products, programmes and APIs that are not models; OpenRouter listings with no located vendor page.
Refresh
Rows are added as vendors publish through October 31, each marked with the date it was added. Corrections are made in place with a dated note. The month closes with a count of releases by type.
Next step

Check the row’s date before you quote it

Every row names the page it was read from and the day it was read. Use the vendor date, not the listing date, and treat the price and latency claims as the vendor’s until a third party measures them.

Agentic AI implementation

Decide which new model earns a slot in your stack

Digital Applied evaluates specialist models against the work they would take off your frontier bill, with a test set from your own tasks.

Task-level test setsCost per taskFallback design
Before you switch

Four checks

  • →The vendor date, not the listing
  • →Price per task, not per token
  • →A test on your own inputs
  • →A second source for the slot
Questions and answers

Practical questions

On October 1: Cloudflare’s Clef and Clef-flash, Amazon’s Strands Decider 2B, Microsoft AI’s MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, and Tavus’s Griffin-Lite preview. On October 2: Bilibili’s Index-Translate-35B-A3B-preview.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source