Ask a large cloud provider for a synthetic voice in English, Mandarin, Spanish, Japanese or German and you will get dozens, most of them good enough for a phone line. Ask for one in a language spoken by a few million well-off people and the answer changes: one or two voices, a reading style that sounds like a news bulletin, no dialect, no handling of the English words that real speech is full of. Sometimes there is nothing at all.
That gap is where custom text-to-speech training earns its keep. This is our map of it. Two caveats before the list: catalogues change every quarter, so check by listening to the current stock voices before you decide anything; and "well served" is about conversational quality on a phone line, not about whether a language code appears in a pricing table.
Why the long tail is long
A production-grade voice needs tens of hours of clean, consistent, licensed studio speech from a single speaker, a text corpus in the same variety of the language, and people who can judge whether the result sounds right. For a language with a hundred million speakers that investment is obvious. For a language with half a million speakers, or a dialect that is never written down, it is not, however wealthy the speakers are. So the big catalogues cover the standard, written form of the big languages and stop.
The consequences show up in three ways:
- Absent. No voice at all. The language is simply not in the catalogue.
- Standard-only. A voice exists for the official written form, which nobody speaks on the phone. Arabic is the largest case: stock voices read Modern Standard Arabic, and callers in Riyadh, Dubai or Cairo speak something else.
- Thin. One voice, one style, one gender, no control over pronunciation of names, and English words mangled.
The markets where this bites, and pays
| Market | Language or variety | What the stock voices usually offer | Why it matters commercially |
|---|---|---|---|
| Saudi Arabia, UAE, Qatar, Kuwait, Bahrain, Oman | Gulf and Najdi Arabic, with English mixed in | Modern Standard Arabic voices, read-style; dialect and code-switching weak | High-income markets with phone-heavy banking, government and hospitality |
| Switzerland | Swiss German dialects | Standard German with a Swiss accent at best; the dialects have no written standard and no stock voices | Banking, insurance and healthcare that speak dialect to customers |
| Luxembourg | Luxembourgish | Absent or nearly so | A financial centre where the national language has almost no speech tooling |
| Norway | Nynorsk and regional dialects | Bokmål voices exist; Nynorsk and dialects thin or absent | Public sector and regional services that must use them |
| Iceland, Faroe Islands, Greenland | Icelandic, Faroese, Greenlandic | Icelandic has a few voices after public investment; Faroese and Greenlandic largely absent | Small, wealthy, and obliged to serve citizens in the language |
| Malta, Ireland, Wales, Basque Country | Maltese, Irish, Welsh, Basque | A voice or two, usually read-style, mostly from public programmes | Official-language obligations in prosperous places |
| Singapore | Singapore English and code-switched Malay, Mandarin, Tamil | Generic English voices; the local variety and the switching are not there | One of the richest markets on earth, phone-heavy, multilingual by default |
| Hong Kong and Taiwan | Hong Kong Cantonese with English, Taiwanese Hokkien | Cantonese voices exist; the mixing is weak; Hokkien largely absent | Finance, retail and public services that speak it |
| Israel | Hebrew | Voices exist; choice and conversational quality lag the big languages | Technology and finance with a demanding domestic market |
| Austria, Belgium, Canada | Austrian German, Flemish, Québec French | Voices exist but few; national accents flattened toward the larger standard | Customers notice when a brand sounds foreign |
If your market is on this list, the stock voice will be the thing your customers remember, and not fondly.
What custom training changes
A custom voice is built from your own recordings: a speaker you choose, recorded with consent and the rights settled, in the variety of the language your customers actually speak, with a pronunciation dictionary for your names and products. The result is yours, runs where you need it, and does not change when a vendor updates a catalogue.
The same applies, with more force, to recognition. If the stock voice cannot speak Gulf Arabic, the stock recogniser cannot hear it either, and a voice agent needs both. Our deepest work is in Gulf and Najdi Arabic for exactly that reason. The method is the same for any of the rows above: collect, transcribe, train, measure on your own audio.
Before you commission anything
Three checks that take an afternoon:
- Listen to the current stock voices in your language on a phone, not on headphones. Play them to three people who speak the variety. If they laugh, you have your answer.
- Count the English words in a hundred real customer sentences. If it is more than a handful per sentence, code-switching is your main requirement, and it is the one stock voices handle worst.
- Find your speaker. A custom voice is only as good as the person and the recording. Settle the consent and the rights before the first session, in writing.
Then bring us the recordings. Twenty minutes is enough to say what can be done with them, and what it would take.
