Test a multilingual voice agent with conversations that change language inside a turn and across turns. A caller may use a second language for a product name, explain a problem in their preferred language and then quote an English reference number. The agent must preserve the task and important details while choosing an appropriate language for its reply. A supported-language list cannot establish that behavior.
- 01Test actual switchesSeparate multilingual availability from mixed-language conversation behavior.
- 02Preserve names and valuesA language change must not rewrite identifiers, amounts or dates.
- 03Score each stageListening, interpretation, speaking and tool actions can fail independently.
- 04Use fluent reviewNaturalness and meaning need reviewers who understand the tested languages.
01 — Language policyDefine the conversation behavior you want
Decide how the assistant should choose its reply language before evaluating it. A business might follow an explicit customer preference, continue in the language established at the start of the call or ask when the preference becomes unclear. There is no single correct rule for every service, but an unstated rule makes consistent evaluation impossible.
In a hypothetical call, a customer speaks Spanish, uses an English product name and continues in Spanish. The expected behavior might be to keep answering in Spanish while preserving the product name. That is different from a customer explicitly asking to continue in English. The test needs both cases so the agent does not treat every borrowed term as a request to switch.
Keep the policy understandable to the caller. If a requested language is unavailable in the deployed workflow, the agent should say so and offer the supported alternative or a handoff. It should not claim fluency because one underlying component lists the language. The voice model comparison provides product context; this guide focuses on the behavior of the assembled conversation.
Specify how long a preference persists. A caller might choose a language for the whole session or ask for one phrase to be explained in another language. The distinction should be represented in the conversation state, not inferred afresh from every utterance. Include the preference in a human handoff so the next participant can continue naturally without asking the customer to repeat a choice the system already confirmed.
Mixed-language phrase
A caller borrows a term or quotes wording without requesting a new reply language.
Explicit language change
A caller asks the assistant to continue in another supported language.
Return to the prior language
A quoted phrase ends and the caller resumes the established conversation.
02 — Test materialBuild paired scripts with the same business meaning
Create a small set of ordinary tasks, such as locating an order, asking about a policy or changing a draft appointment request. Write a single-language version and a mixed-language version with the same intended outcome. Keep the business facts stable so a difference in results can be traced to the language pattern rather than a harder task.
Use synthetic names, addresses and references. Include a product name that should remain unchanged, a number that needs confirmation and a phrase whose meaning depends on context. A fluent reviewer should check that the mixed-language script is something a real speaker could naturally say. Artificially alternating languages on every word can be a useful stress case, but it should not stand in for ordinary calls.
Keep the expected interpretation separate from the exact wording. Several transcripts or spoken replies may preserve the meaning correctly. The test should accept legitimate variants while identifying changes to the task, named entity or requested action. Otherwise, it becomes a string-matching exercise that penalizes natural language without measuring operational correctness.
Keep routine and adversarial examples separate in the results. A natural mixed-language request measures everyday service quality; an intentionally confusing switch tests the boundary at which the system should ask for help. Both are useful, but combining them without labels makes a failure rate difficult to interpret. The team needs to know whether the agent struggles with ordinary calls or behaves appropriately when the input itself is genuinely ambiguous.
The scripts in this method are proposed evaluation material. No language pair or provider has been benchmarked for this article.
03 — Listening pathSeparate recognition from understanding
When a call goes wrong, inspect what each stage received. In a transcription-based pipeline, compare the audio with the transcript and then inspect the agent's interpretation. A correct transcript can still lead to the wrong action. An imperfect transcript may preserve enough meaning for a safe clarification. Those outcomes need different repairs.
A native audio model may not expose a transcript that represents every part of its internal processing. If the system provides a transcript for review, do not automatically treat it as a complete trace of the model's reasoning. Inspect the actual tool arguments and response as well. The observable business outcome remains the final acceptance surface.
ElevenLabs' model documentation, checked October 11, 2026, lists language support across speech-generation and recognition models. That helps select candidates for a trial, but it does not prove reliable code-switching in a particular call. The deployed combination of models, configuration and transport needs its own evaluation.
Preserve timing in the evidence where it affects meaning. A pause may separate a quoted phrase from the next instruction, and a correction may arrive before or after the agent commits a value. A transcript flattened into one paragraph loses that sequence. Keep enough audio or timestamped event information to distinguish a recognition error from a turn-taking error, while limiting retention to the evaluation material the team is authorized to store.
- Compare the recording with any available transcript.
- Inspect extracted names, numbers and task intent.
- Check the actual tool arguments before judging the reply.
04 — Stable factsProtect names and numbers during a switch
A language change should not translate an account identifier, alter a product code or silently reinterpret an ambiguous date. Decide which values are literal and which need locale-aware interpretation. When the meaning is uncertain, confirmation is more useful than a fluent guess. The customer should hear the value that the system is about to use.
For a hypothetical booking call, a customer gives a day and month in one language, then repeats a number in another. The test should verify the resolved date rather than merely checking that both languages appeared in the transcript. Keep the expected date explicit in the test record and require clarification when the utterance permits more than one interpretation.
Our pronunciation test guide covers how the agent speaks names and values. Code-switching adds the question of whether it preserved them while the surrounding language changed. A pronunciation dictionary can improve a spoken form where supported, but it does not establish the correct customer identity or repair a misunderstood instruction.
Test read-back as a business control rather than a decorative repetition. The agent should confirm the ambiguous value in a form the caller can meaningfully correct. Repeating an unfamiliar identifier too quickly may not help, even when every character is technically spoken. Define how the conversation proceeds when the caller disagrees, and inspect whether the corrected value replaces the prior one in the pending action rather than merely appearing in the transcript.
A name can be recognized correctly and pronounced badly, or pronounced naturally after being attached to the wrong record. Test both meaning and speech.
05 — Preference memoryTest reply language across the full conversation
A short demonstration often shows a successful switch and stops there. Continue the conversation. Ask a follow-up, introduce a quoted phrase and return to the original language. The agent should follow the defined policy without drifting because the most recent product name happened to belong to another language.
Include an explicit correction such as a request to speak more slowly or to continue in a different supported language. Check whether the agent changes only what was requested. A language preference should not reset the booking details or cause the system to ask again for information it already confirmed. Conversely, a corrected business detail must update the task even if the reply language stays the same.
The conversation-flow documentation describes configurable behavior in a particular agent platform. Use such controls as implementation inputs, then test the observable conversation. A setting's label is not evidence that interruption handling, language preference and tool state remain aligned in your workflow.
Include a brief unrelated utterance from another speaker only if it is relevant to the intended environment and recorded with appropriate consent. The agent should not automatically change the customer's language preference because background speech uses another language. This tests the boundary between the active conversation and incidental audio. Do not claim universal speaker separation from one sample; record the conditions and the behavior the particular workflow demonstrated.
| Scenario | Expected check | Distinct failure |
|---|---|---|
| Borrowed product name | Reply follows the established preference | Unwanted language switch |
| Explicit switch request | Next reply uses the requested supported language | Preference ignored |
| Quoted foreign phrase | Meaning preserved without task reset | Quote treated as a new instruction |
| Return to prior language | Behavior follows the stated policy | Conversation drifts unpredictably |
06 — Evaluation qualityUse natural recordings and fluent reviewers
Start with controlled recordings so candidate systems receive the same input. Include speakers and speaking conditions relevant to the intended service, with their permission. A clean studio script is useful for diagnosis, but it does not represent every accent, hesitation or telephone connection the system may encounter.
Have reviewers score meaning preservation, reply-language choice, intelligibility and business action separately. Ask them to identify the phrase that caused a failure and the consequence for the caller. An overall naturalness score can hide a wrong number or a missed request to change language. Keep recordings and review notes under the same access controls as other test material.
Do not infer that an accent predicts proficiency or customer value. The practical purpose is to find where the system misunderstands speech and improve the service or handoff. If coverage is limited to a small set of speakers, say so in the evaluation record. A successful sample does not establish reliable behavior for an entire language community.
Give reviewers a shared rubric with examples of acceptable variation. One reviewer may prefer a regional expression while another considers a different expression natural. Separate that preference from a meaning-changing error, an unintelligible phrase or an inappropriate language choice. Record disagreements and resolve consequential ones with source context rather than averaging away a failure that changed the customer's instruction or the action the agent was about to take.
- Use consented, representative recordings and synthetic business details.
- Keep fluent reviewers focused on meaning as well as naturalness.
- Report tested language pairs and conditions without generalizing beyond them.
07 — Business outcomeCheck the action after the conversation sounds right
A fluent reply can conceal an incorrect tool call. Inspect the destination state for a proposed appointment, lookup or update in an isolated test environment. Confirm that the agent selected the intended record and preserved the values from the call. Do not accept a spoken confirmation as proof that the operation happened correctly.
Include a language switch while a tool is pending and a correction after the first result. Those cases test whether the agent can keep conversation state separate from action state. Our background voice-work guide explains why stopping speech does not by itself cancel a background action. A multilingual correction needs the same explicit reconciliation.
When a failure cannot be resolved confidently, evaluate the handoff. The next person should receive the customer's preferred language, confirmed facts and unresolved detail without a fabricated summary. The translation QA guide provides a related principle: preserve business meaning rather than treating fluency as sufficient acceptance.
Exercise a failed tool result in each tested language pattern. The agent must explain that the action did not complete without translating an error into a confident confirmation. A multilingual success demo says little about this path because the system never has to distinguish a rejected request from a completed task. Preserve the tool response, spoken explanation and destination state together so the reviewer can see whether the failure was communicated accurately.
A correct spoken sentence and an incorrect destination record are a failed test. Review both surfaces together.
08 — Rollout scopeRelease the language pairs you have actually tested
Define the initial supported conversation patterns from the evidence you collected. A team may be ready for one language pair in a narrow appointment workflow while still investigating another pair or a noisy-call condition. Make that scope operationally clear so the agent can offer a legitimate alternative when the request exceeds it.
Keep failed examples for regression testing when a model, voice or conversation setting changes. Improvements in one component can alter pacing, language detection or the timing of tool calls elsewhere. Repeat the cases that previously failed rather than relying on a fresh demonstration that happens to avoid them.
Our AI transformation service helps teams turn language support into a tested workflow with clear handoffs and observable outcomes. The useful result is a voice agent that preserves the caller's task through a language change, and knows when it needs help, not a claim to handle every language combination because a model card lists many languages.
Give support staff a concise statement of what the system can currently handle and how to report a difficult call. Useful reports identify the language pair, task, misunderstood phrase and observed consequence without making assumptions about the caller. That feedback can extend the evaluation set deliberately. It should not become a reason to broaden the public support claim before the new cases have been reviewed and the relevant failure repaired.
- Document the tested language pairs, tasks and conditions.
- Keep a fallback for unsupported or uncertain conversations.
- Retest difficult cases when any speech or agent component changes.
Test the switch and the task together
Use paired scenarios with stable business facts, inspect the listening and action stages, and ask fluent reviewers to assess the conversation. Continue past the first successful language switch.
Approve only the patterns the assembled workflow has demonstrated. Language availability is a starting point for evaluation, not its conclusion.