An hour of recorded speech costs between ten cents and about a dollar to turn into text at list price, depending on the provider and whether it is live. The same hour costs under a cent to feed into the cheapest model that listens to it directly, and between about eleven cents and several dollars to feed in as video, depending on how many pixels the model is shown each second. Those three sentences describe three different tasks, and this page keeps them in three tables for that reason.
Every figure is a list price read from the provider's own pricing page or documentation on September 18, 2026, converted to one hour with the arithmetic shown, using the tokens-per-second rate each provider states. Where a provider's own per-hour figure exists, it is checked against the conversion and the match or mismatch is noted. Commercial speech vendors are named, not linked. Output tokens, the transcript or summary the model writes, are extra in the second and third tables and are billed at each model's output rate.
- 01Batch transcription clusters at $0.15 to $0.36 an hour.Microsoft's preview MAI-Transcribe-2 is the outlier at $0.10 through December 31, 2026. Live transcription costs more, from $0.15 to $1.02, and Deepgram's live rate is marked promotional.
- 02Feeding audio to a multimodal model can cost less than transcribing it.Qwen3.8-Omni-Flash bills seven tokens a second at $0.15 per million, which is $0.0038 an hour of input. Gemini's Flash-Lite models land at three to six cents. That buys understanding, not a verbatim transcript, and output is extra.
- 03Video price is a resolution setting.At one frame a second, Gemini 3.8 Flash costs $0.275 an hour at low resolution and $0.84 at high. Qwen's $0.20 for 720p is the vendor's chart figure; its documentation supports a range of $0.16 to $0.32.
- 04Two prices carry dates.Gemini 3.8 Flash input doubles from $0.75 to $1.50 per million tokens on January 1, 2027. Microsoft's $0.10 an hour is stated only through December 31, 2026.
01 — DefinitionsThree tasks, three price shapes
Speech-to-text transcription turns audio into words and is priced per minute or per hour of audio; the output is the transcript and is included. Audio understanding means sending the recording to a multimodal language model and asking it a question, such as "summarise this call" or "who agreed to what"; it is priced per audio token, at a tokens-per-second rate the provider sets, and the answer is billed separately as output. Video understanding is the same idea with frames added, and its price depends on the resolution and frame rate the model is shown. A transcription row and an understanding row cannot be ranked against each other, because one includes the output and the other does not, and because they produce different things.
The per-hour conversions use one formula for token-priced rows: tokens per second, times 3,600, times the price per million, divided by a million. Per-minute prices are multiplied by 60 and per-second prices by 3,600. Where we had to assume a rate, the row says "est." and the methodology explains why.
02 — Table 1Transcription per hour
Pre-recorded, pay-as-you-go, standard tier, in ascending order of cost per hour. Volume discounts, enterprise tiers and add-ons are noted only where the pricing page shows them.
| Model | List price | Per hour | Note |
|---|---|---|---|
| Microsoft MAI-Transcribe-2 | $0.10 / hr | $0.10 | Public preview; price stated through Dec 31, 2026 |
| Alibaba Cloud Qwen3-ASR (qwen3-asr-flash, intl.) | $0.000035 / s | $0.126 | 0.000035 × 3,600 |
| AssemblyAI Universal-2 | $0.15 / hr | $0.15 | |
| Mistral Voxtral Mini Transcribe 2 | $0.003 / min | $0.18 | 0.003 × 60 |
| Meta Muse Voice Transcribe 1.0 | $0.18 / hr | $0.18 | Streaming priced the same |
| OpenAI gpt-4o-mini-transcribe | $0.003 / min (est.) | $0.18 | Token-billed; per-minute figure is OpenAI's estimate |
| Google Cloud Speech-to-Text V2, dynamic batch | $0.003 / min | $0.18 | Lower-urgency queue; Standard tier |
| AssemblyAI Universal-3.5 Pro | $0.21 / hr | $0.21 | Diarization add-on $0.02 / hr |
| ElevenLabs Scribe v2 | $0.22 / hr | $0.22 | |
| Mistral Voxtral Small | $0.004 / min | $0.24 | Also an audio-chat model; one rate on the card |
| Deepgram Nova-3, monolingual | $0.0043 / min | $0.258 | Multilingual $0.0052 / min = $0.312 |
| OpenAI gpt-transcribe | $0.0045 / min | $0.27 | |
| Google Gemini 3.5 Transcribe | $2 / M in + $12 / M out | ≈ $0.31 | 25 tok/s in, 175 tok/min out; Google's own blend is ~$0.30 |
| OpenAI Whisper | $0.006 / min | $0.36 | |
| OpenAI gpt-4o-transcribe and -diarize | $0.006 / min (est.) | $0.36 | |
| Google Cloud Speech-to-Text V2, Standard | $0.016 / min | $0.96 | First 500k min/month; falls to $0.24 / hr above 2M min |
Live transcription is a separate sub-table because most providers price it differently and one of the rates below is marked promotional.
| Model | List price | Per hour | Note |
|---|---|---|---|
| AssemblyAI Universal-Streaming | $0.15 / hr | $0.15 | |
| Meta Muse Voice Transcribe 1.0 | $0.18 / hr | $0.18 | Same as batch |
| Deepgram Nova-3, monolingual | $0.0048 / min | $0.288 | Marked promotional; regular $0.0077 / min = $0.462 |
| ElevenLabs Scribe v2 Realtime | $0.39 / hr | $0.39 | |
| AssemblyAI Universal-3.5 Pro Realtime | $0.45 / hr | $0.45 | |
| Google Gemini 3.5 Transcribe Live | $3.50 / M in + $21 / M out | ≈ $0.54 | Same token rates as batch; Google's blend ~$0.009 / min |
| OpenAI gpt-live-transcribe | $0.017 / min | $1.02 | Also listed as gpt-realtime-whisper |
For a closer look at OpenAI's transcription models against each other, see our transcription model comparison; for the case where the right answer is no per-hour bill at all, see our guide to self-hosted transcription.
03 — Table 2Audio understanding per hour
Input only. The tokens-per-second column is what the provider documents for audio input; the per-hour figure is that rate times 3,600 times the input price. OpenAI documents its ten-tokens-a- second rate for the Realtime API, and we apply it to the gpt-audio models as an estimate, marked as such.
| Model | Audio tokens / s | Input price / M | Per hour of input |
|---|---|---|---|
| Qwen3.8-Omni-Flash (international) | 7 | $0.15 | $0.0038 |
| Gemini 3.5 Flash-Lite | 32 | $0.30 | $0.035 |
| Gemini 3.1 Flash-Lite | 32 | $0.50 (audio) | $0.058 |
| Gemini 3.8 Flash | 32 | $0.75 to Dec 31, 2026 | $0.086 |
| Gemini 3.1 Pro Preview | 32 | $2.00 (≤200k prompt) | $0.23 |
| Qwen3.5-Omni-Plus (international) | 7 | $11 (audio) | $0.277 |
| OpenAI gpt-audio-mini / gpt-realtime-2.1-mini | 10 (est.) | $10 | $0.36 (est.) |
| OpenAI gpt-audio / gpt-audio-1.5 / gpt-realtime-2.1 | 10 (est.) | $32 | $1.15 (est.) |
Cost of one hour of audio input, multimodal models
Provider pricing pages and token documentation, read September 18, 2026. OpenAI rows use the Realtime API token rate as an estimate. Input only; output billed separately.The Qwen row is the one behind the claim in our Qwen3.8-Omni-Flash post. On this method the model's hour of audio input costs $0.0038 against $0.277 for its predecessor Qwen3.5-Omni-Plus, a reduction of 98.6%, which agrees with Qwen's "more than 98%"; Qwen's chart prints the two figures as under $0.01 and $0.28. Gemini 3.8 Flash at $0.086 also agrees with the $0.09 on Qwen's chart, so the chart's audio panel reproduces from public documentation.
04 — Table 3Video understanding per hour
One hour of video with its audio track, input only, at one frame per second unless stated. Google bills video by frame at 70 tokens for its low and medium settings and 280 for high, plus 32 audio tokens a second. Alibaba Cloud's estimator resizes a 720p frame to 594 tokens at 32-pixel tiles, plus 7 audio tokens a second. Change the frame rate or the resolution and every figure below changes with it.
| Model | Assumption | Per hour | Arithmetic or note |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | Low resolution, 1 fps | $0.11 | 102 tok/s × 3,600 × $0.30 / M |
| Gemini 3.1 Flash-Lite | Low resolution, 1 fps | $0.12 | Video at $0.25 / M, audio at $0.50 / M |
| Qwen3.8-Omni-Flash | 720p, 1 fps (Qwen's method) | $0.20 | Vendor-run chart figure; docs-based estimate $0.16 to $0.32 |
| Gemini 3.8 Flash | Low resolution, 1 fps | $0.275 | 102 tok/s × 3,600 × $0.75 / M |
| Gemini 3.1 Flash-Lite | High resolution, 1 fps | $0.31 | 280 + 32 tok/s at split rates |
| Gemini 3.5 Flash-Lite | High resolution, 1 fps | $0.34 | 312 tok/s × 3,600 × $0.30 / M |
| Gemini 3.8 Live (streaming video) | Google's per-minute rate | $0.42 | $0.002 / min video + $0.005 / min audio, × 60 |
| Gemini 3.1 Pro Preview | Low resolution, 1 fps, chunked ≤200k tokens | $0.73 | One unchunked hour exceeds 200k and bills at $4 / M: $1.47 |
| Gemini 3.8 Flash | High resolution, 1 fps | $0.84 | Matches Qwen's chart, which used high resolution |
| Gemini 3.1 Pro Preview | High resolution, 1 fps, chunked | $2.25 | Unchunked: $4.49 |
| Qwen3.5-Omni-Plus | 720p, 1 fps | $3.27 | 594 video + 7 audio tok/s; matches Qwen's chart exactly |
Applying the same token method that reproduces Qwen's own Qwen3.5-Omni-Plus figure to the cent, 594 video tokens plus 7 audio tokens a second at $0.15 per million, gives $0.32 an hour for Qwen3.8-Omni-Flash. If the estimator's pairing of two frames per temporal token applies, it gives $0.16. Qwen's chart says $0.20, which sits between the two and which the public documentation does not let us reproduce exactly. The row prints Qwen's number, labelled, with the range beside it.
05 — FindingsFindings
Four things stand out once the conversions are done. First, batch transcription has converged: ten of the sixteen rows sit between $0.15 and $0.27 an hour, and the cheapest published rate, Microsoft's $0.10, is a preview price with an end date. The outlier at the top, Google Cloud's $0.96 for the first half-million minutes, falls to $0.24 at volume, so the spread is mostly a question of tier.
Second, listening is now cheaper than transcribing, if listening is what you need. An hour of audio into Qwen3.8-Omni-Flash costs less than a cent and into Gemini's Flash-Lite models a few cents, against fifteen to thirty-six cents to transcribe it. The catch is in the word "input": the model still has to write something, and a full verbatim transcript as output would cost more than the input did. For summaries, answers and action items, though, the arithmetic favours the multimodal route by an order of magnitude.
Third, video cost is a setting you choose. The same hour into Gemini 3.8 Flash costs $0.275 at low resolution and $0.84 at high, and a single unchunked hour into Gemini 3.1 Pro crosses the 200,000-token threshold that doubles its rate. Anyone quoting a video-understanding price without the resolution and frame rate is quoting half a number.
Fourth, two prices on this page are dated and one is promotional. Gemini 3.8 Flash doubles on January 1, 2027, Microsoft's transcription rate is stated only to December 31, 2026, and Deepgram's streaming rate is marked as a limited-time promotion. A budget built on this table should carry those three notes with it. If you would like the table applied to your own volumes, with the output side estimated for your workload, our AI transformation service does that as a short engagement.
06 — How to read thisMethodology
A census of list prices with every conversion shown, not a benchmark of quality. Three task groups, never ranked against each other.
- Inclusion rule
- A row needs a price on the provider's own pricing page, documentation or model card, a stated unit, and, for token-priced models, a documented tokens-per-second rate for the modality. Prices reported only by resellers, aggregators or press are excluded; an aggregator listing was used once, to cross-check, and where it disagreed with the provider the provider's figure stands.
- Sources
- OpenAI API pricing and voice cost guide; Google Gemini API pricing (page dated September 16, 2026), tokens and media-resolution documentation; Google Cloud Speech-to-Text pricing; Microsoft Tech Community launch post for MAI-Transcribe-2 (September 3, 2026); Meta developer pricing and model pages; Mistral model cards for Voxtral; Alibaba Cloud Model Studio pricing (international and Chinese Mainland, updated September 18, 2026) and Qwen-Omni documentation; Qwen's Qwen3.8-Omni-Flash post and chart; Deepgram, AssemblyAI and ElevenLabs pricing pages. All read September 18, 2026.
- Conversions
- Token rows: tokens per second × 3,600 × price per million ÷ 1,000,000. Per-minute prices × 60; per-second prices × 3,600. Tokens per second as documented: Gemini audio 32, Gemini Transcribe 25 in and 175 text tokens per minute out, OpenAI 10 for user audio under the Realtime API, Qwen-Omni 7. Video: Gemini 70 tokens per frame at low or medium resolution and 280 at high, plus audio, at 1 fps; Qwen 720p resized to 1056 × 576 pixels, which is 594 tokens per frame at 32-pixel tiles, plus 7 audio tokens a second.
- Currency
- USD list prices for international deployments. Qwen3.8-Omni-Flash is also listed for Chinese Mainland at 0.8 yuan input, 0.1 yuan cache-hit and 2.7 yuan output per million tokens, which gives 0.020 yuan per hour of audio input; no exchange conversion was applied. Qwen's own footnote converts at 6.7191 yuan to the dollar.
- As-of date
- All pages read September 18, 2026. Only Google's pricing page, the Microsoft post and Alibaba Cloud's pages carry their own dates; the others are undated, so the collection date is the only date that applies to them.
- Known limitations
- OpenAI's ten-tokens-a-second audio rate is documented for the Realtime API and applied to gpt-audio as an estimate. Google's tokens page gives two video rates; we used the media-resolution table, which matches the $0.84 on Qwen's own chart. Google Cloud's Chirp 3 price is inferred from the Standard tier. Microsoft's price comes from its launch post only, because the Azure pricing page loads prices by script and the retail prices API returned no meter. Mistral's pricing page rendered no Voxtral figures, so its model cards were used. Meta's Muse Spark audio and video prices are not published and are omitted. Qwen's $0.20 video figure is vendor-run and not exactly reproducible from documentation.
07 — Next stepCents an hour, if you know which task you are buying
Price your workload from the token counts the API returns, not from this table
Use the tables to shortlist and to catch a vendor quoting the wrong task. Then run one hour of your own media through the two or three candidates, read the input and output token counts the API returns, and multiply by the list price yourself. That number includes the output side this page cannot estimate for you, and it is the one to put in the budget. When a listed price changes, this page will be updated with a new as-of date.