AI DevelopmentMethodology6 min readPublished September 26, 2026

180 routes · 8 output types · one catalog snapshot · the AI models that do not answer in words

AI Models That Make Images, Video and Speech: 180 Counted

One model catalog lists 180 AI routes that output images, video, speech or other non-text results. Counts by type, every video route, and the price gaps.

DA
Digital Applied Team
Research and practical guidance
Non-text routes180 of 625
ObservedSep 25, 2026

Most talk about AI models is about chat. But in the snapshot of OpenRouter’s public model catalog we took on September 25, 2026, 180 of the 625 routes produce something other than text: images, video, speech, transcripts and the number lists that power search. They come from 35 vendors, and they are being added faster than chat models.

A route here means one model offered through the catalog’s single API, so a model with a batch or free variant counts more than once. The count below is what a developer or marketer can reach through one gateway, not a census of every model in the world. It also shows a practical gap: for 39 of these routes, including every video model, the catalog prints no usable price.

Key takeaways
  1. 01
    180 of 625 routes, 29 percent, output something other than text.Image leads with 57 routes, then embeddings with 37 and video with 29.
  2. 02
    Non-text routes are arriving faster.75 of the 180 were listed in the 90 days to September 25, against 129 of 445 text-only routes: 42 percent against 29.
  3. 03
    No price is shown for 39 non-text routes.All 29 video routes, six rerank routes, two music routes and two image routes list zero in every price field.
  4. 04
    A listing is not proof a model is available.The snapshot still listed Sora 2 Pro a day after OpenAI’s scheduled removal of the Sora 2 API.

01 — The countMore than a quarter of the catalog does not talk

Share of catalog
Routes with a non-text output
29%180 of 625

Counted from every route in the all-modality catalog on September 25. The other 445 output text only.

Snapshot
Vendors
Vendor namespaces behind the 180 routes
35plus 2 routers

OpenAI has the most with 23, then Google with 20, Recraft with 16 and Qwen with 10.

By vendor
No price shown
Every price field reads zero
39routes

Not counting six routes marked free or OpenRouter’s two automatic routers, whose price depends on the model they pick.

Price gap

The pace is the clearest signal. Non-text listings ran at 29 in July and 27 in August, and 18 more arrived in the first 25 days of September. For comparison with language models, our census of context and output limits covers the text side of the market.

02 — The tableRoutes by output type

Some output types need a word of explanation. Embeddings turn text into lists of numbers so software can find similar content; rerank models sort search results by relevance. Transcription turns speech into text and speech turns text into a voice. Audio covers two Google music models and two OpenAI voice-chat models. Decisions is a newer type that returns a structured choice instead of prose; our post on TypeSafe’s Jev explains it.

Source: OpenRouter models API, all output modalities, observed September 25, 2026 at 08:01 UTC. “Added” means listed on or after June 27, 2026.
OutputRoutesAdded in 90 daysNo price shown
Image57232
Embeddings3770
Video291329
Transcription24140
Speech20130
Rerank736
Audio402
Decisions220
All non-text1807539
Text only (comparison)445129—

Non-text routes by output type

OpenRouter models API, observed September 25, 2026
Image
57
Embeddings
37
Video
29
Transcription
24
Speech
20
Rerank
7
Audio
4
Decisions
2

Fifteen routes produce text as well as another output, mostly Google and OpenAI image models that can reply in words and pictures. Each counts once under its non-text type. Among speech routes, 16 of 20 list their voices: Deepgram’s Aura-2 lists 90, Kokoro 54 and MiniMax’s Speech 2.8 models 45 each, while Microsoft’s MAI Voice 2 lists four. Our ranking of text-to-speech models compares their quality and vendor prices.

03 — The listEvery video route in the catalog

Most of the 29 video routes turn a prompt or a still image into a clip. Nine also accept video, for editing, extending or upscaling footage, and six also accept audio as an input.

Source: OpenRouter models API, observed September 25, 2026. “Listed” is the catalog’s own date, not the vendor’s release date.
Catalog IDListedAccepts
black-forest-labs/flux-video-editSep 10, 2026Text, video
minimax/hailuo-3-maxSep 2, 2026Text, image
alibaba/wan-3.0-primeAug 27, 2026Text, image
alibaba/wan-3.0Aug 24, 2026Text, image
heygen/avatar-ivAug 24, 2026Text, image, audio
black-forest-labs/flux-video-upscaleAug 19, 2026Text, video
bytedance/seedance-2.0-miniAug 12, 2026Text, image, audio, video
bytedance/seedance-2.5Aug 7, 2026Text, image, audio, video
black-forest-labs/flux-3-videoAug 4, 2026Text, image, video
minimax/hailuo-3Jul 29, 2026Text, image, audio, video
runway/aleph-2Jul 29, 2026Text, image, video
runway/gen-4.5Jul 29, 2026Text, image
x-ai/grok-imagine-video-1.5Jul 20, 2026Text, image
alibaba/happyhorse-1.1Jun 24, 2026Text, image
alibaba/happyhorse-1.0Jun 24, 2026Text, image
x-ai/grok-imagine-videoMay 18, 2026Text, image
kwaivgi/kling-v3.0-proApr 29, 2026Text, image
kwaivgi/kling-v3.0-stdApr 29, 2026Text, image
google/veo-3.1-fastApr 24, 2026Text, image
google/veo-3.1-liteApr 23, 2026Text, image
kwaivgi/kling-video-o1Apr 20, 2026Text, image
minimax/hailuo-2.3Apr 20, 2026Text, image
alibaba/wan-2.7Apr 15, 2026Text, image
bytedance/seedance-2.0Apr 15, 2026Text, image, audio, video
bytedance/seedance-2.0-fastApr 15, 2026Text, image, audio, video
alibaba/wan-2.6Mar 28, 2026Text, image
bytedance/seedance-1-5-proMar 23, 2026Text, image
openai/sora-2-proMar 23, 2026Text, image
google/veo-3.1Mar 23, 2026Text, image

One row shows why a catalog is a starting point, not an authority. OpenAI’s deprecations page scheduled the removal of its Videos API and the Sora 2 models, including sora-2-pro, for September 24. Our snapshot a day later still listed openai/sora-2-pro. The catalog also shows an expiry date of November 11, 2026 on bytedance/seedance-1-5-pro, the only non-text route carrying one.

04 — The catchWhat the price fields cannot tell you

For chat models, a catalog price is a cost per million tokens and can be compared directly. For other outputs it cannot. Every video route lists zero in every price field, so the catalog gives no way to budget a clip. Image routes mostly carry their cost in image fields rather than the usual input and output fields: 37 of 57 list zero for both of those.

Transcription shows the unit problem most sharply. Its raw input prices run from 0.00000125 to 0.36 across 24 routes, a spread that reflects different billing units, such as tokens, seconds or minutes, rather than a real price gap. The catalog does not say which unit each route uses.

Zero does not mean free

A zero in a catalog price field means the catalog has no price in that field, not that the model costs nothing. Six routes in this set are explicitly marked free. For the rest, read the vendor’s own pricing page before you plan spend, and run a small paid test to see what the invoice actually counts.

05 — MethodologyHow we counted

Methodology

One snapshot, one rule per column, no inference from model names.

What was collected
Every route in OpenRouter’s models API with all output modalities requested: 625 records, each with its listed input and output types, listing date, price fields, voices and expiry date.
As-of date
September 25, 2026 at 08:01 UTC. Routes added or removed after that moment are not reflected. OpenAI’s deprecations page was read on September 29, 2026.
Classification
A route counts as non-text if any listed output is not text. Output type is the catalog’s own label. “Added in 90 days” uses the catalog’s listing date, on or after June 27, 2026. “No price shown” means every price field is zero, excluding routes marked free and the two automatic routers.
Counting rules
Batch, free and alias variants count as separate routes, as the catalog lists them. Vendors are counted by catalog namespace, excluding OpenRouter’s own routers and counting the ~typesafe alias under typesafe. No route listed more than one non-text output type.
Known limitations
One gateway’s catalog, not the whole market: models sold only direct by their vendors are missing. Listing dates are not release dates, and price fields were recorded raw, without converting units.
Refresh
Refreshed in place from each new catalog snapshot, with counts and the video list updated and the as-of date changed.

06 — ConclusionChoosing a non-text model starts where the catalog stops

Budgeting video generation
Ignore catalog prices, which are all zero. Price a test clip on the vendor’s own terms.
Vendor page
Choosing a voice model
Check the voice list and languages, then compare quality and per-minute cost outside the catalog.
Test voices
Comparing transcription prices
Convert every price to cost per audio hour first. Raw fields use different units.
Normalise
Picking an embedding model
Text prices are listed per token, as for chat models, so they compare more easily. Test retrieval quality on your own documents.
Compare
What to do next

Use the catalog to find candidates, then confirm price, units and availability on each vendor’s own page before you build

Image, video and voice models are now a large and fast-growing part of what one API key can reach, but the catalog describes them less completely than chat models. For embeddings, our embedding cost calculator does the conversion for you. If you want these models built into a content or product workflow with real costs attached, our AI transformation team can help.

Digital Applied

Image, video and voice AI, priced before you build.

We choose and test generation models for your workflow, confirm prices and units on each vendor’s own terms, and wire them into production with cost tracking.

Model selectionCost testsProduction pipelines
Your next project

Generation models that fit the budget

  • →Vendor prices confirmed
  • →Units normalised for comparison
  • →Availability checked before build
Questions and answers

The questions we get about non-text AI models

OpenRouter's catalog listed 29 video routes from 10 vendors on September 25, 2026. Vendors also sell models directly that are not in that catalog, so the full market is larger.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Eleven v4 and v4 Turbo: What Changed for Voice Agents

ElevenLabs' Eleven v4 tops Artificial Analysis' voice arena and v4 Turbo claims ~100ms. The two latency figures, the vendor tests and the v4 price.

September 28, 2026 · 6 minRead
AI Development

Which AI Models Still Let You Set Temperature? 15 Checked

Vendor docs show 10 of 15 current AI models reject a custom temperature or say to drop it. Each rule, where a catalog disagrees, and what to use instead.

September 27, 2026 · 5 minRead
AI Development

An Open Model That Makes Editable, Layered Designs

Ming-Image-0.1-Design makes UI, poster and infographic images with transparent backgrounds; its Layer variant splits a flat design into editable layers. MIT.

September 25, 2026 · 6 minRead
AI Development

Gemini 2.5 Flash Image Retires October 2 on the API: Act Now

The Gemini API shuts down gemini-2.5-flash-image on October 2, 2026; Vertex says March 15, 2027. Which date applies, three replacements priced, and a test plan.

September 25, 2026 · 4 minRead
AI Development

Google Intelligent Eyewear: Gemini AI Glasses Fall 2026

Google announces Gemini-powered smart glasses with Samsung, Gentle Monster, and Warby Parker at I/O. Audio glasses ship fall 2026; display tier TBD.

May 20, 2026 · 18 minRead
AI Development

AI Search Agents Compared: Google, Perplexity, ChatGPT

Google's always-on information agents, Perplexity Pro, and ChatGPT Search compared. Which AI search agent delivers the best research results in 2026?

May 20, 2026 · 14 minRead