AI DevelopmentCost Playbook11 min readPublished October 10, 2026

The cheapest unit may be measuring a different job

Voice AI Pricing: Compare Characters, Tokens and Minutes

Compare the cost of delivering your own scripts and calls, with the provider’s billing units kept intact.

DA
Digital Applied Team
Research and practical guidance
PublishedOctober 10, 2026
Read time11 min
SourcesPrimary documentation

Voice AI prices become comparable when you price the same finished work. A character rate, an audio-token rate and a per-minute rate are not competing quotes until you know what each meter includes. Build the estimate from your actual script or call pattern, then keep generation, listening, reasoning and transport charges separate. This is a costing method, not a new ranking of voice models.

Key takeaways
  1. 01
    Keep the native unitRecord exactly what the provider bills before attempting a conversion.
  2. 02
    Use a shared workloadThe same script or call scenario makes a comparison meaningful.
  3. 03
    Count rejected outputRetries and unused speech still belong in the cost of accepted work.
  4. 04
    Reconcile the invoiceA planning estimate should be checked against actual metered usage.

01 — Workload boundaryStart by naming the voice job

A recorded product explanation and a live appointment assistant can both use speech synthesis, but their bills describe different activities. The recording starts with a fixed script and ends with an approved audio file. The assistant listens, waits, reasons, speaks and may call external tools. Comparing their headline minute prices hides that difference before the arithmetic even starts.

Write a compact workload description. For a hypothetical recording project, specify the approved text, languages, voices and required output format. For a hypothetical telephone assistant, describe the customer turns, approximate response length, expected waiting periods and tasks it may complete. These are planning inputs, not measured usage or a promise about how long real customers will speak.

Keep speech generation apart from audio understanding. Our audio and video processing cost reference covers the cost of processing existing media. This guide concerns comparing the meters inside a voice workflow. A low transcription price cannot establish that a complete conversational agent is cheap, because it accounts for only one part of the service.

For a repeatable brief, specify what is outside the job too. A recorded announcement does not include a conversational model merely because the same provider offers one. A live assistant may use a separate model for reasoning even when its speech sounds like a single product. Listing the actual components prevents a buyer from comparing a narrowly priced synthesis endpoint with a managed conversation service and mistaking the difference in scope for a difference in efficiency.

Fixed content
An approved recording
Generation workload

Compare the same script, voice requirements and accepted audio deliverable.

One artifact
Interactive service
A completed conversation
Multiple meters

Include listening, speaking, reasoning, transport and waiting behavior.

One outcome

02 — Meter definitionsKeep four billing units separate

Characters describe a text quantity under a provider's counting rules. Tokens describe the units used by a particular model or tokenizer. Audio duration describes time in a recording or stream. Connected minutes describe elapsed session time under the service's contract. Each can be valid, but none is automatically a substitute for another.

A token is especially easy to misread. Text input tokens, audio input tokens and audio output tokens may have different rates and counting rules. The word token alone does not identify the meter. Store the complete usage field name and the endpoint next to the rate; otherwise, a spreadsheet can apply an input rate to output usage without producing an obvious mathematical error.

ElevenLabs' model documentation, checked October 11, 2026, describes different model families and character-credit treatment. That is a reason to check the selected model and plan together, not a reason to assume a credit always equals one character. A direct account, a bundled agent product and a gateway may expose different commercial units for related capabilities.

Create separate columns for the quantity unit and the rate denominator. “Audio output tokens” and “per million audio output tokens” should both remain visible. If the usage export reports seconds but the contract prices minutes, record that explicit duration conversion and any rounding rule. This is different from estimating speaking time from written characters: seconds and minutes measure the same physical quantity, while characters and duration describe different properties of the workload.

Billing distinctions for comparison design. The table contains no vendor price quotes or universal conversion factors.
MeterWhat to recordDo not assume
CharactersProvider-counted submitted textA fixed speaking duration
TokensExact text or audio usage fieldAll token types cost the same
Audio durationInput or generated audio timeOnly accepted audio is charged
Connected timeContract-defined session durationSilence is excluded

03 — Conversion trapDo not turn characters into minutes by guesswork

The same written text can produce different durations because of language, punctuation, pauses, speaking style and speed. A list of reference numbers also behaves differently from a paragraph of ordinary prose. A rough conversion may be useful for an early budget, but it must remain an assumption with a range rather than a fact about the provider's billing.

Use the approved script when comparing fixed recordings. Generate an evaluation sample through each candidate endpoint, retain its reported billable usage and measure the resulting file's duration. The cost per delivered minute can then be derived from that sample. Repeat with materially different content before using the relationship across an entire library. This article proposes that procedure; it does not report a completed provider experiment.

For live calls, scripted tests are useful but cannot fully reproduce caller behavior. A hurried speaker, a long pause or a correction changes the mix of listening and speaking. Keep an initial estimate separate from observed production usage. If the provider supplies an explicit audio-token conversion, record its scope and date rather than transferring it to another model with a similar name.

A useful sensitivity check changes one assumption at a time. Keep the approved text fixed and compare a slower delivery setting, then a different language, then a script with more numbers. Record whether the accepted duration and metered quantity change. The resulting range belongs to those observations. Do not turn the most favorable sample into a universal conversion that understates the expected bill for every future recording or every caller in the service.

Practical check

A conversion factor belongs to a specific model, workload and counting rule. Record where it came from, and label estimates as estimates.

04 — Cost arithmeticBuild the bill from independent line items

The safest calculation multiplies each measured quantity by its own applicable rate, then adds the results. For a hypothetical character-priced service charging $12 per million characters, 50,000 billable characters would cost $0.60. For a separate hypothetical duration-priced service charging $0.08 per generated minute, 15 minutes would cost $1.20. These invented rates illustrate arithmetic only; the two workloads are not asserted to be equivalent.

A live agent can add several more terms: speech recognition, the language model, synthesized responses, telephone transport and separately charged tools. A bundled price may already include some of them. Mark included items explicitly so the same charge is not added twice. Conversely, do not treat a blank spreadsheet cell as evidence that an item is included or free.

Subscription commitments deserve their own row. Compare the incremental cost of one more job and the total cost of the expected workload separately. Included credits that expire unused still affect the effective cost of the work you actually complete. Our model human-review cost guide extends the same reasoning to the time spent checking output after the machine bill is known.

For procurement, keep a second view that compares total monthly scenarios using the same expected mix of jobs. A plan with a commitment can be economical at one volume and expensive at another, but the break point depends on the actual included units and overage schedule. If a sales quote bundles several services, ask for enough scope detail to reconstruct the comparison. A lower total with an unspecified inclusion is not yet a comparable offer.

  • Record the exact meter, rate, currency, endpoint and observation date.
  • Separate included usage from overage and minimum commitments.
  • Keep taxes, transport and review outside a quoted model-only total.

05 — Output yieldCount speech that never reaches the customer

Generated audio can be rejected for pronunciation, pacing, an incorrect detail or an unsuitable tone. It may also be generated ahead of playback and discarded when a caller interrupts. Whether each event is billable depends on the service, so preserve usage records rather than assuming that unheard audio is free. The customer-facing artifact is only part of the production history.

For an illustrative recording workflow, divide all generation charges for the job by the accepted audio duration. Do not divide by the combined duration of every attempt: that makes repeated failures look productive. Keep the original meter alongside the derived figure so another person can reconstruct the result. If no recording was accepted, report an unsuccessful job rather than a misleading zero cost per accepted minute.

The same principle applies to calls, although completion needs a business definition. A call that connects and ends is not necessarily a successfully handled request. Track the machine bill for abandoned, escalated and completed calls, then compare those groups without hiding the difficult cases. A model that speaks less may save money, or it may be failing to explain the next step; the outcome record tells you which.

Tag rejected attempts by cause so the team knows what kind of repair would improve yield. A pronunciation failure may need a text preparation change; an incorrect business fact needs a better source or approval step; an unsuitable voice may require a different creative choice. Combining those failures into one retry count loses the operational explanation. Keep the quality category beside the cost rather than asking a single price-per-minute figure to describe both.

Useful denominator

Use accepted audio for a recording project and a clearly defined completed task for an agent. Keep failure costs in the numerator.

06 — Evaluation designUse a small comparison with a stable script

Choose representative material before opening the provider dashboards. Include ordinary explanations, business names, numbers and at least one correction. For multilingual work, use approved text in each language rather than assuming one language's duration and tokenization will stand in for another. The test should resemble the production job closely enough that a difference has an operational meaning.

Keep output requirements fixed: file format, voice suitability, intelligibility and any latency constraint that actually matters. A beautifully expressive recording is not interchangeable with a response that must arrive during a live conversation. Review quality before calculating the price of acceptable output, then record why a candidate failed instead of silently removing it from the table.

The pronunciation test guide provides a companion acceptance method for names and numbers. Do not let the costing exercise become a quality ranking unsupported by listening tests. An invoice proves usage and charges; it does not prove that a customer could understand the result or that the agent carried out the requested action.

Use a blind listening order where practical so a familiar provider name does not decide the quality result in advance. Give reviewers the intended audience and script requirements, not just a request to choose their favorite voice. A clear but restrained reading may be more suitable for an account instruction than a highly expressive performance. The accepted-output comparison should reflect the job's requirements rather than personal preference alone.

  • Use the same business content and acceptance criteria.
  • Retain request usage, audio duration and rejection reason together.
  • Repeat materially different languages and interaction patterns separately.

07 — Invoice checkReconcile the estimate with actual usage

After a controlled trial, compare the application's request records with the provider's usage report. Allow for documented reporting delays and rounding rules, but investigate unexplained differences. Request identifiers help connect a surprising charge to retries, longer output or a separate endpoint. Without them, the team is left arguing about totals that cannot be traced to work.

Check whether the application is re-synthesizing identical approved messages unnecessarily. Reusing a permitted, current audio asset can change the workload, but caching needs an explicit key that includes the text, voice, settings and language. A stale greeting containing the wrong opening hours is not a useful saving. Confirm the service's terms before treating generated assets as freely reusable across contexts.

Set an operational response to unexpected spend. It might pause optional generation, route a request for review or stop a looping call; the right action depends on the job. A budget alert alone does not control a runaway workflow. Our AI transformation service connects these cost records with task design so a saving does not come from removing the checks that make the agent useful.

Investigate duplicate requests before blaming a provider's meter. A client reconnect, an application timeout or an eager retry policy may submit the same text several times. Preserve the application's attempt identifier and the provider's request identifier together. If a request was accepted but the response was lost, the usage report may be the first evidence that work occurred. Resolving that uncertainty can prevent the application from buying the same output again.

Practical check

An estimate is ready for broader use when its assumptions explain the observed bill. A close total with unexplained line-item differences is not yet a reliable model.

08 — Purchase decisionChoose the service on accepted work

The useful comparison is a short explanation of what the workload costs, what quality it delivers and where the estimate remains uncertain. Preserve the native billing units even when you present a derived cost per call or recording. That lets the buyer change volume assumptions without rebuilding the research from a headline price.

A recording-heavy team may find a character-based plan predictable once its scripts are stable. A conversational team may prefer a bundle whose scope is clear, even if one component appears more expensive in isolation. Neither conclusion follows from the billing unit itself. It follows from the task mix, included services, review needs and measured output yield.

Revisit the calculation when the model, voice settings, languages or commercial plan changes. A price cut on one meter can coincide with a workflow change that consumes more of it. Keep the decision attached to a dated workload description and a repeatable acceptance process; that is what turns a pricing comparison into a budget a team can actually operate.

Give the final comparison a named owner and a clear trigger for rechecking it, such as a contract renewal or a change in the deployed endpoint. A dated record is useful only if a future operator can tell which assumptions still apply. Keep the original test script and accepted files alongside the calculation. That makes the next review an update to an existing workload model rather than another search for a convenient headline rate.

  • Choose a workload that represents the intended use.
  • Keep the original usage meters visible in the final comparison.
  • Explain uncertainties before expanding the budget.
Your next step

Price one complete voice workload

Start with the script or conversation your business needs, preserve each billing unit and include rejected attempts. Compare the cost of acceptable work only after the output has passed the same review.

A clear meter-by-meter record will remain useful when prices change. A guessed characters-to-minutes conversion usually will not.

Put the method to work

Build a workflow your team can verify

Digital Applied helps teams turn a promising AI capability into a clear operating process, with useful evaluations, review points and a practical path to production.

Workflow designPractical evaluationsClear ownership
Work with us

From trial to useful work

  • →Define the task and its acceptance criteria
  • →Connect the right information and tools
  • →Review failures before expanding access
FAQ · Practical implementation

Questions before you start

Only after defining what the minute measures and how the figure was derived. Generated audio, input audio and connected call time are different quantities.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue exploring

Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source