AI DevelopmentPricing Tracker8 min readPublished September 18, 2026

42 rows · 9 providers · every conversion shown · transcription is not understanding

What an Hour of Audio or Video Costs to Process With AI

List prices for one hour of audio or video across nine providers, from OpenAI and Google to Alibaba and the speech vendors, with every conversion shown.

DA
Digital Applied Team
Research and practical guidance
Editorial dateSeptember 18, 2026
Data as ofSeptember 18, 2026

An hour of recorded speech costs between ten cents and about a dollar to turn into text at list price, depending on the provider and whether it is live. The same hour costs under a cent to feed into the cheapest model that listens to it directly, and between about eleven cents and several dollars to feed in as video, depending on how many pixels the model is shown each second. Those three sentences describe three different tasks, and this page keeps them in three tables for that reason.

Every figure is a list price read from the provider's own pricing page or documentation on September 18, 2026, converted to one hour with the arithmetic shown, using the tokens-per-second rate each provider states. Where a provider's own per-hour figure exists, it is checked against the conversion and the match or mismatch is noted. Commercial speech vendors are named, not linked. Output tokens, the transcript or summary the model writes, are extra in the second and third tables and are billed at each model's output rate.

Key takeaways
  1. 01
    Batch transcription clusters at $0.15 to $0.36 an hour.Microsoft's preview MAI-Transcribe-2 is the outlier at $0.10 through December 31, 2026. Live transcription costs more, from $0.15 to $1.02, and Deepgram's live rate is marked promotional.
  2. 02
    Feeding audio to a multimodal model can cost less than transcribing it.Qwen3.8-Omni-Flash bills seven tokens a second at $0.15 per million, which is $0.0038 an hour of input. Gemini's Flash-Lite models land at three to six cents. That buys understanding, not a verbatim transcript, and output is extra.
  3. 03
    Video price is a resolution setting.At one frame a second, Gemini 3.8 Flash costs $0.275 an hour at low resolution and $0.84 at high. Qwen's $0.20 for 720p is the vendor's chart figure; its documentation supports a range of $0.16 to $0.32.
  4. 04
    Two prices carry dates.Gemini 3.8 Flash input doubles from $0.75 to $1.50 per million tokens on January 1, 2027. Microsoft's $0.10 an hour is stated only through December 31, 2026.

01DefinitionsThree tasks, three price shapes

Speech-to-text transcription turns audio into words and is priced per minute or per hour of audio; the output is the transcript and is included. Audio understanding means sending the recording to a multimodal language model and asking it a question, such as "summarise this call" or "who agreed to what"; it is priced per audio token, at a tokens-per-second rate the provider sets, and the answer is billed separately as output. Video understanding is the same idea with frames added, and its price depends on the resolution and frame rate the model is shown. A transcription row and an understanding row cannot be ranked against each other, because one includes the output and the other does not, and because they produce different things.

The per-hour conversions use one formula for token-priced rows: tokens per second, times 3,600, times the price per million, divided by a million. Per-minute prices are multiplied by 60 and per-second prices by 3,600. Where we had to assume a rate, the row says "est." and the methodology explains why.

02Table 1Transcription per hour

Pre-recorded, pay-as-you-go, standard tier, in ascending order of cost per hour. Volume discounts, enterprise tiers and add-ons are noted only where the pricing page shows them.

Batch speech-to-text list prices from each provider's pricing page or model card, read September 18, 2026. OpenAI's per-minute figures for its token-billed models are OpenAI's own estimates.
ModelList pricePer hourNote
Microsoft MAI-Transcribe-2$0.10 / hr$0.10Public preview; price stated through Dec 31, 2026
Alibaba Cloud Qwen3-ASR (qwen3-asr-flash, intl.)$0.000035 / s$0.1260.000035 × 3,600
AssemblyAI Universal-2$0.15 / hr$0.15
Mistral Voxtral Mini Transcribe 2$0.003 / min$0.180.003 × 60
Meta Muse Voice Transcribe 1.0$0.18 / hr$0.18Streaming priced the same
OpenAI gpt-4o-mini-transcribe$0.003 / min (est.)$0.18Token-billed; per-minute figure is OpenAI's estimate
Google Cloud Speech-to-Text V2, dynamic batch$0.003 / min$0.18Lower-urgency queue; Standard tier
AssemblyAI Universal-3.5 Pro$0.21 / hr$0.21Diarization add-on $0.02 / hr
ElevenLabs Scribe v2$0.22 / hr$0.22
Mistral Voxtral Small$0.004 / min$0.24Also an audio-chat model; one rate on the card
Deepgram Nova-3, monolingual$0.0043 / min$0.258Multilingual $0.0052 / min = $0.312
OpenAI gpt-transcribe$0.0045 / min$0.27
Google Gemini 3.5 Transcribe$2 / M in + $12 / M out≈ $0.3125 tok/s in, 175 tok/min out; Google's own blend is ~$0.30
OpenAI Whisper$0.006 / min$0.36
OpenAI gpt-4o-transcribe and -diarize$0.006 / min (est.)$0.36
Google Cloud Speech-to-Text V2, Standard$0.016 / min$0.96First 500k min/month; falls to $0.24 / hr above 2M min

Live transcription is a separate sub-table because most providers price it differently and one of the rates below is marked promotional.

Real-time speech-to-text list prices, read September 18, 2026. Deepgram marks its streaming rate as a limited-time promotion with no end date shown.
ModelList pricePer hourNote
AssemblyAI Universal-Streaming$0.15 / hr$0.15
Meta Muse Voice Transcribe 1.0$0.18 / hr$0.18Same as batch
Deepgram Nova-3, monolingual$0.0048 / min$0.288Marked promotional; regular $0.0077 / min = $0.462
ElevenLabs Scribe v2 Realtime$0.39 / hr$0.39
AssemblyAI Universal-3.5 Pro Realtime$0.45 / hr$0.45
Google Gemini 3.5 Transcribe Live$3.50 / M in + $21 / M out≈ $0.54Same token rates as batch; Google's blend ~$0.009 / min
OpenAI gpt-live-transcribe$0.017 / min$1.02Also listed as gpt-realtime-whisper

For a closer look at OpenAI's transcription models against each other, see our transcription model comparison; for the case where the right answer is no per-hour bill at all, see our guide to self-hosted transcription.

03Table 2Audio understanding per hour

Input only. The tokens-per-second column is what the provider documents for audio input; the per-hour figure is that rate times 3,600 times the input price. OpenAI documents its ten-tokens-a- second rate for the Realtime API, and we apply it to the gpt-audio models as an estimate, marked as such.

Audio input to multimodal models, list prices per million audio tokens, read September 18, 2026. Output tokens are extra. Gemini 3.8 Live is omitted here because Google prices it per minute ($0.005 / min audio, $0.30 / hr).
ModelAudio tokens / sInput price / MPer hour of input
Qwen3.8-Omni-Flash (international)7$0.15$0.0038
Gemini 3.5 Flash-Lite32$0.30$0.035
Gemini 3.1 Flash-Lite32$0.50 (audio)$0.058
Gemini 3.8 Flash32$0.75 to Dec 31, 2026$0.086
Gemini 3.1 Pro Preview32$2.00 (≤200k prompt)$0.23
Qwen3.5-Omni-Plus (international)7$11 (audio)$0.277
OpenAI gpt-audio-mini / gpt-realtime-2.1-mini10 (est.)$10$0.36 (est.)
OpenAI gpt-audio / gpt-audio-1.5 / gpt-realtime-2.110 (est.)$32$1.15 (est.)

Cost of one hour of audio input, multimodal models

Provider pricing pages and token documentation, read September 18, 2026. OpenAI rows use the Realtime API token rate as an estimate. Input only; output billed separately.
Qwen3.8-Omni-Flash7 tok/s × $0.15/M
$0.0038
Gemini 3.5 Flash-Lite32 tok/s × $0.30/M
$0.035
Gemini 3.1 Flash-Lite32 tok/s × $0.50/M
$0.058
Gemini 3.8 Flash32 tok/s × $0.75/M
$0.086
Gemini 3.1 Pro Preview32 tok/s × $2.00/M
$0.23
Qwen3.5-Omni-Plus7 tok/s × $11/M
$0.277
OpenAI gpt-audio-mini10 tok/s × $10/M, est.
$0.36
OpenAI gpt-audio10 tok/s × $32/M, est.
$1.15

The Qwen row is the one behind the claim in our Qwen3.8-Omni-Flash post. On this method the model's hour of audio input costs $0.0038 against $0.277 for its predecessor Qwen3.5-Omni-Plus, a reduction of 98.6%, which agrees with Qwen's "more than 98%"; Qwen's chart prints the two figures as under $0.01 and $0.28. Gemini 3.8 Flash at $0.086 also agrees with the $0.09 on Qwen's chart, so the chart's audio panel reproduces from public documentation.

04Table 3Video understanding per hour

One hour of video with its audio track, input only, at one frame per second unless stated. Google bills video by frame at 70 tokens for its low and medium settings and 280 for high, plus 32 audio tokens a second. Alibaba Cloud's estimator resizes a 720p frame to 594 tokens at 32-pixel tiles, plus 7 audio tokens a second. Change the frame rate or the resolution and every figure below changes with it.

Video input list prices with the stated resolution and frame-rate assumption, read September 18, 2026. Output tokens are extra. Gemini 3.8 Flash input prices double on January 1, 2027.
ModelAssumptionPer hourArithmetic or note
Gemini 3.5 Flash-LiteLow resolution, 1 fps$0.11102 tok/s × 3,600 × $0.30 / M
Gemini 3.1 Flash-LiteLow resolution, 1 fps$0.12Video at $0.25 / M, audio at $0.50 / M
Qwen3.8-Omni-Flash720p, 1 fps (Qwen's method)$0.20Vendor-run chart figure; docs-based estimate $0.16 to $0.32
Gemini 3.8 FlashLow resolution, 1 fps$0.275102 tok/s × 3,600 × $0.75 / M
Gemini 3.1 Flash-LiteHigh resolution, 1 fps$0.31280 + 32 tok/s at split rates
Gemini 3.5 Flash-LiteHigh resolution, 1 fps$0.34312 tok/s × 3,600 × $0.30 / M
Gemini 3.8 Live (streaming video)Google's per-minute rate$0.42$0.002 / min video + $0.005 / min audio, × 60
Gemini 3.1 Pro PreviewLow resolution, 1 fps, chunked ≤200k tokens$0.73One unchunked hour exceeds 200k and bills at $4 / M: $1.47
Gemini 3.8 FlashHigh resolution, 1 fps$0.84Matches Qwen's chart, which used high resolution
Gemini 3.1 Pro PreviewHigh resolution, 1 fps, chunked$2.25Unchunked: $4.49
Qwen3.5-Omni-Plus720p, 1 fps$3.27594 video + 7 audio tok/s; matches Qwen's chart exactly
Why Qwen's video figure is marked vendor-run

Applying the same token method that reproduces Qwen's own Qwen3.5-Omni-Plus figure to the cent, 594 video tokens plus 7 audio tokens a second at $0.15 per million, gives $0.32 an hour for Qwen3.8-Omni-Flash. If the estimator's pairing of two frames per temporal token applies, it gives $0.16. Qwen's chart says $0.20, which sits between the two and which the public documentation does not let us reproduce exactly. The row prints Qwen's number, labelled, with the range beside it.

05FindingsFindings

Four things stand out once the conversions are done. First, batch transcription has converged: ten of the sixteen rows sit between $0.15 and $0.27 an hour, and the cheapest published rate, Microsoft's $0.10, is a preview price with an end date. The outlier at the top, Google Cloud's $0.96 for the first half-million minutes, falls to $0.24 at volume, so the spread is mostly a question of tier.

Second, listening is now cheaper than transcribing, if listening is what you need. An hour of audio into Qwen3.8-Omni-Flash costs less than a cent and into Gemini's Flash-Lite models a few cents, against fifteen to thirty-six cents to transcribe it. The catch is in the word "input": the model still has to write something, and a full verbatim transcript as output would cost more than the input did. For summaries, answers and action items, though, the arithmetic favours the multimodal route by an order of magnitude.

Third, video cost is a setting you choose. The same hour into Gemini 3.8 Flash costs $0.275 at low resolution and $0.84 at high, and a single unchunked hour into Gemini 3.1 Pro crosses the 200,000-token threshold that doubles its rate. Anyone quoting a video-understanding price without the resolution and frame rate is quoting half a number.

Fourth, two prices on this page are dated and one is promotional. Gemini 3.8 Flash doubles on January 1, 2027, Microsoft's transcription rate is stated only to December 31, 2026, and Deepgram's streaming rate is marked as a limited-time promotion. A budget built on this table should carry those three notes with it. If you would like the table applied to your own volumes, with the output side estimated for your workload, our AI transformation service does that as a short engagement.

06How to read thisMethodology

Methodology

A census of list prices with every conversion shown, not a benchmark of quality. Three task groups, never ranked against each other.

Inclusion rule
A row needs a price on the provider's own pricing page, documentation or model card, a stated unit, and, for token-priced models, a documented tokens-per-second rate for the modality. Prices reported only by resellers, aggregators or press are excluded; an aggregator listing was used once, to cross-check, and where it disagreed with the provider the provider's figure stands.
Sources
OpenAI API pricing and voice cost guide; Google Gemini API pricing (page dated September 16, 2026), tokens and media-resolution documentation; Google Cloud Speech-to-Text pricing; Microsoft Tech Community launch post for MAI-Transcribe-2 (September 3, 2026); Meta developer pricing and model pages; Mistral model cards for Voxtral; Alibaba Cloud Model Studio pricing (international and Chinese Mainland, updated September 18, 2026) and Qwen-Omni documentation; Qwen's Qwen3.8-Omni-Flash post and chart; Deepgram, AssemblyAI and ElevenLabs pricing pages. All read September 18, 2026.
Conversions
Token rows: tokens per second × 3,600 × price per million ÷ 1,000,000. Per-minute prices × 60; per-second prices × 3,600. Tokens per second as documented: Gemini audio 32, Gemini Transcribe 25 in and 175 text tokens per minute out, OpenAI 10 for user audio under the Realtime API, Qwen-Omni 7. Video: Gemini 70 tokens per frame at low or medium resolution and 280 at high, plus audio, at 1 fps; Qwen 720p resized to 1056 × 576 pixels, which is 594 tokens per frame at 32-pixel tiles, plus 7 audio tokens a second.
Currency
USD list prices for international deployments. Qwen3.8-Omni-Flash is also listed for Chinese Mainland at 0.8 yuan input, 0.1 yuan cache-hit and 2.7 yuan output per million tokens, which gives 0.020 yuan per hour of audio input; no exchange conversion was applied. Qwen's own footnote converts at 6.7191 yuan to the dollar.
As-of date
All pages read September 18, 2026. Only Google's pricing page, the Microsoft post and Alibaba Cloud's pages carry their own dates; the others are undated, so the collection date is the only date that applies to them.
Known limitations
OpenAI's ten-tokens-a-second audio rate is documented for the Realtime API and applied to gpt-audio as an estimate. Google's tokens page gives two video rates; we used the media-resolution table, which matches the $0.84 on Qwen's own chart. Google Cloud's Chirp 3 price is inferred from the Standard tier. Microsoft's price comes from its launch post only, because the Azure pricing page loads prices by script and the retail prices API returned no meter. Mistral's pricing page rendered no Voxtral figures, so its model cards were used. Meta's Muse Spark audio and video prices are not published and are omitted. Qwen's $0.20 video figure is vendor-run and not exactly reproducible from documentation.

07Next stepCents an hour, if you know which task you are buying

Put it into practice

Price your workload from the token counts the API returns, not from this table

Use the tables to shortlist and to catch a vendor quoting the wrong task. Then run one hour of your own media through the two or three candidates, read the input and output token counts the API returns, and multiply by the list price yourself. That number includes the output side this page cannot estimate for you, and it is the one to put in the budget. When a listed price changes, this page will be updated with a new as-of date.

Digital Applied

Budget the audio or video product before you build it.

We run your own recordings through the candidate models, log the real token counts, and hand you a cost-per-hour figure with the output side included, then build on the model that wins.

Real token accountingThree-model bake-offDated price notes
Your next project

Start with one hour

  • Pick one representative recording
  • Run it through three candidates
  • Multiply returned tokens by list price
Questions and answers

Applying this post

Because they are different products with different price shapes. Transcription is priced per minute or hour and the transcript is included. Audio understanding is priced per audio token for the input, and whatever the model writes back is billed separately as output. A per-hour transcription price and a per-hour audio-input price are not the same kind of number, so this page never ranks them together.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Qwen3.8-Omni-Flash: Cheaper Audio and Video AI Agents

Qwen's September 18 model takes audio and video natively with 1M context and, by Qwen's own method, cuts per-hour audio cost 98% versus its predecessor.

September 18, 2026 · 9 minRead
AI Development

Where Your Code Goes: A Census of Agent Data Terms

Seventeen coding agents, six questions each, vendor documentation only. What the published terms say on retention and training, and which cells stayed open.

August 17, 2026 · 18 minRead
AI Development

Kimi Code With K3: Setup, Plans and Cache Discipline

Running Kimi K3 in Kimi Code: the Moderato plan unlocks 256K context, Allegretto the full 1M, and cache discipline — not knobs — controls your real cost.

July 17, 2026 · 10 minRead
AI Development

Multimodal AI for Marketing: Applications and Strategies

Leverage multimodal AI in marketing: GPT-5.2, Gemini 3 Pro, Claude for image, video, and audio content. Real use cases and implementation strategies.

January 18, 2026 · 13 minRead
AI Development

Deleting AI Agent Memory: Where Stored Copies Survive

Deleting AI agent memory takes more than clearing a chat. Map stored copies, retrieval indexes and backups, then verify what your system can still recover.

September 4, 2026 · 6 minRead
AI Development

Preview, Beta, GA: What Vendors Said vs What Coverage Said

Thirty-six AI vendor announcements from 17-22 August 2026, each scored on the vendor's own status word against the word its coverage used, where located.

August 22, 2026 · 27 minRead