Test voice-agent pronunciation with the words and numbers your business cannot afford to make ambiguous: customer names, place names, prices, dates and reference codes. Keep the intended text beside the generated audio, then check the result through the same channel a caller will use. A pleasing voice is not evidence that these details are being spoken correctly.
The method below is a proposed quality check, not a benchmark we have already run. Use synthetic records for the first pass, obtain approval for any real recordings you later use, and have appropriate listeners decide the acceptable pronunciations. Regional variation and personal preference should not be mistaken for errors.
- 01Find the failing layerIncorrect source text, ambiguous formatting and faulty speech output need different fixes.
- 02Agree on acceptable readingsNames and regional pronunciations can have more than one valid form.
- 03Verify dictionary supportA dictionary entry is useful only if the chosen model and endpoint apply it.
- 04Test the delivered audioA clean generated file is not the same as intelligible speech on the customer's actual call path.
01 — Test selectionChoose the details that carry business meaning
Begin with information that changes what the listener understands or does. A price, an appointment date and an order reference deserve more attention than a decorative opening sentence. The aim is not to find a voice that pronounces every possible word perfectly. It is to establish that the system communicates the important details of your own service clearly and can recover when the listener is unsure.
Ask the people who handle customer conversations which words are regularly repeated, spelled out or corrected. Their list may include local place names, product abbreviations and surnames that a generic demonstration never encounters. Turn those categories into synthetic test cases so the early evaluation does not depend on customer recordings. Keep the test small enough that a person can listen carefully to every result.
For each case, write the intended meaning and the acceptable spoken forms. A person's preferred name pronunciation should come from that person or an authorized source, not a generated guess. If two regional forms are both acceptable for your audience, record that explicitly. Otherwise a reviewer may mark a valid variation as a failure and a later reviewer may reverse the decision without changing the audio.
| Case type | What to preserve | What the listener checks |
|---|---|---|
| Person or place | Approved name and pronunciation | The intended identity is recognizable |
| Money | Amount and currency | The listener hears the correct value |
| Date and time | Calendar meaning and time zone | The appointment is unambiguous |
| Reference code | Character sequence | The identifier can be repeated accurately |
| Abbreviation | Intended expansion or letter reading | The phrase has the business meaning intended |
| Address | Street, number and locality | The destination is not confused with another |
02 — Failure diagnosisSeparate the sentence from the sound
Keep the text sent to speech generation with the returned audio. If the language model changed the price before the voice spoke it, a pronunciation dictionary is the wrong repair. If the text contains an ambiguous date, the system needs a normalization rule before synthesis. If the text is correct and explicit but the audio is wrong, the speech layer becomes the focus. Without this record, every problem is blamed on the voice.
Consider a synthetic booking reference containing both a letter O and a zero. The source record may be correct, the generated response may preserve it, and the voice may still read it in a way that a caller cannot distinguish. The practical repair could be a deliberate spoken format with pauses and explicit character names. That is different from changing the stored identifier or teaching a dictionary a new word.
Review the complete chain once before optimizing any one layer. The transcript-verification guide explains the related distinction between what was heard and what a transcript records. For speech output, preserve the authoritative value, the display text, the speech text and the final audio as separate artifacts. A correction to one should not silently rewrite the others.
The wrong value was written
Check the source record and language-model output before touching pronunciation settings.
The right value sounded wrong
Inspect normalization, dictionary support and the chosen voice on the actual endpoint.
03 — NormalizationMake numbers unambiguous before synthesis
Numbers often need a business-specific spoken form. An identifier should usually retain its character sequence, while a price should communicate an amount and currency. A date should be expanded enough to avoid a locale-dependent interpretation. The application should know which kind of value it is speaking rather than asking a general language model to infer the meaning from punctuation.
For example, the string 12/10 can represent different calendar dates under different conventions. The useful repair is to speak an explicit month and day based on the authoritative record, not to choose a more expressive voice. Similarly, a decimal amount without a currency may be pronounced clearly while still leaving the customer unsure what they will pay. Pronunciation quality cannot compensate for missing semantics.
Keep transformations narrow and reversible. A rule that expands an abbreviation in a company name may be wrong when the same letters occur inside an account code. A rule that removes punctuation to improve speech may destroy the structure of an address or reference. Test the normalizer with both the case it is intended to fix and a nearby case it must leave unchanged. The resulting spoken text should remain traceable to the original business value.
- Classify the value before selecting a spoken format.
- Expand ambiguous dates, currencies and abbreviations deliberately.
- Keep identifiers intact even when their spoken form adds pauses or labels.
04 — Pronunciation controlsUse dictionaries with model-specific expectations
The W3C Pronunciation Lexicon Specification describes a way to associate written forms with pronunciation information, including phonemes and aliases. A standard format can make a lexicon understandable and portable in principle, but it does not prove that every modern speech endpoint accepts every feature. Support remains something to verify for the actual service and model.
ElevenLabs' pronunciation dictionary documentation, checked October 11, lists phoneme-tag support for eleven_v4, eleven_flash_v2 and eleven_v3, and recommends aliases for other models that skip those tags. That distinction matters operationally: uploading a syntactically valid dictionary can still leave pronunciation unchanged when the selected model does not apply the relevant feature.
Treat a proposed dictionary entry as a hypothesis to test. Generate the word in an ordinary sentence, at the beginning of a reply and near another difficult term. Then ask an appropriate listener whether the result is acceptable. A phonetic transcription produced by another AI system is a draft, not an authority on a person's name. Keep the approved reading, the dictionary version and the audio sample together so the team knows what was actually verified.
Deliberately test one term whose pronunciation should change when the dictionary is enabled. If the audio does not reflect the intended change, investigate support and attachment before expanding the dictionary.
05 — Audio pathListen in context and through the delivery channel
A term can sound clear in isolation and become confusing inside a sentence. Surrounding words affect rhythm, pauses and emphasis, and callers do not normally hear a carefully selected word sample. Test the phrases the agent will actually use, including a short confirmation, an explanation and a correction. Keep those contexts consistent when comparing a change to the dictionary or voice.
Use the same output path as the intended service. A desktop audio file, a browser conversation and a telephone call can differ in processing and playback conditions. The purpose of this check is not to claim that one channel is always worse. It is to discover whether the delivered audio remains intelligible under the conditions your customers will encounter. A pristine internal preview is only one part of that evidence.
Ask listeners to repeat the important value before showing them the expected text. Showing the answer first can make an ambiguous recording seem clearer than it is. After the first listening, compare the repeated value with the intended one and record the type of disagreement. Keep subjective voice preference separate from whether the amount, name or reference was understood. The voice-benchmark guide provides useful context for avoiding a broad quality claim from a narrow test.
- Listen without displaying the expected answer first.
- Test full sentences as well as isolated terms.
- Repeat the check through the intended customer channel.
06 — Scoring decisionsMake the review record useful for a repair
A single pass or fail label is often too coarse to guide the next change. Record whether the spoken content was incorrect, ambiguous, unusual but acceptable, or clear. Add the listener's interpretation and the exact configuration that generated the audio. That record lets the team distinguish a consistent failure from a reviewer preference without pretending that every disagreement has the same consequence.
Choose acceptance rules based on the use of the phrase. An ambiguous account code may need an immediate repair because it prevents the caller from identifying a record. A stylistic preference in a greeting may be a lower priority. Do not average the two into a reassuring overall score that hides the important failure. Report the critical cases separately and retain the unsuccessful audio samples.
For a hypothetical test set, a useful outcome is a list of required repairs and the cases each repair should affect. It is not an invented accuracy percentage or a claim that the selected voice is ready for every accent and language. If the team needs numerical acceptance thresholds, define them before running the test and explain their business rationale. The AI review-cost guide helps connect this evaluation effort to the value of the work being automated.
A friendly-sounding answer with the wrong amount is not rescued by strong ratings for tone. Track meaning-critical failures independently from preferences about delivery.
07 — Regression checksRetest corrections and nearby words
After a repair, rerun the original failure and a small set of neighboring cases. A spelling alias that improves one name can affect another word or change the rhythm of a longer phrase. A normalization rule that helps a phone number can damage a product code. The objective is not merely to demonstrate that the selected example sounds better; it is to check that the change has not moved the error elsewhere.
Keep a stable set of accepted examples for model, voice and dictionary changes. Record all three versions where the service exposes them. A new model may handle the same instruction differently, and a new voice may emphasize a term differently even when the text is unchanged. If an endpoint uses a moving alias, note that limitation instead of claiming perfect reproducibility from a model name alone.
Assign an owner for additions to the pronunciation list. Without ownership, every reported issue can become a one-off prompt adjustment that nobody retests. A modest review process is often enough: collect the term, confirm its intended reading, test the proposed fix and add the accepted example to the regression set. Our AI transformation service uses this kind of narrow operating record to make quality work repeatable rather than dependent on the person who built the first demo.
- Repeat the original failed phrase after the change.
- Check similar words and values for unintended effects.
- Keep an accepted audio sample and the configuration that produced it.
08 — Live recoveryGive the caller a way to correct the agent
Even a carefully tested voice agent needs a recovery path. A caller may use a name that was not in the test set or prefer a pronunciation the business did not anticipate. The agent should be able to ask for clarification, spell a reference, repeat a value in another form or transfer the conversation when the information cannot be confirmed. Repeating the same unclear sentence with more confidence is not recovery.
Make sure a correction changes the right thing. If the caller explains how to pronounce a name, that should not automatically alter the stored spelling. If they correct an amount, the system should verify the underlying record rather than merely speak the new number back as fact. The distinction between presentation and business data remains important throughout the conversation.
A practical launch decision combines the test record with these fallback behaviors. The team should know which important terms passed, which remain unresolved and what the agent does when a listener cannot understand it. That is a more useful definition of readiness than a general statement that the voice sounds natural. It also gives the business a way to improve the system from real feedback without converting every complaint into an uncontrolled configuration change.
When someone supplies their preferred pronunciation, treat it as a preference to confirm and apply in the appropriate scope. Do not insist that a dictionary or model knows better.
Approve the meaning, not just the voice
Build a small test from the names and values your business actually speaks, preserve the intended text and listen through the real delivery path. Diagnose the layer that failed before changing settings.
A good launch record shows what was tested, which readings were accepted, which configuration produced them and how the agent recovers when a caller needs a correction.