Skip to content
Zingaro AI
Blog · 21 September 2026 · 4 min read

Where the stock voices still fail: wealthy markets, under-served languages

Text-to-speech is excellent in about twenty languages and thin everywhere else. Some of the places where it is thin are rich, which is why custom voice training is a business. A map, with the caveats.

By Zingaro AI

  • voice
  • models
  • markets

Ask a large cloud provider for a synthetic voice in English, Mandarin, Spanish, Japanese or German and you will get dozens, most of them good enough for a phone line. Ask for one in a language spoken by a few million well-off people and the answer changes: one or two voices, a reading style that sounds like a news bulletin, no dialect, no handling of the English words that real speech is full of. Sometimes there is nothing at all.

That gap is where custom text-to-speech training earns its keep. This is our map of it. Two caveats before the list: catalogues change every quarter, so check by listening to the current stock voices before you decide anything; and "well served" is about conversational quality on a phone line, not about whether a language code appears in a pricing table.

Why the long tail is long

A production-grade voice needs tens of hours of clean, consistent, licensed studio speech from a single speaker, a text corpus in the same variety of the language, and people who can judge whether the result sounds right. For a language with a hundred million speakers that investment is obvious. For a language with half a million speakers, or a dialect that is never written down, it is not, however wealthy the speakers are. So the big catalogues cover the standard, written form of the big languages and stop.

The consequences show up in three ways:

  • Absent. No voice at all. The language is simply not in the catalogue.
  • Standard-only. A voice exists for the official written form, which nobody speaks on the phone. Arabic is the largest case: stock voices read Modern Standard Arabic, and callers in Riyadh, Dubai or Cairo speak something else.
  • Thin. One voice, one style, one gender, no control over pronunciation of names, and English words mangled.

The markets where this bites, and pays

MarketLanguage or varietyWhat the stock voices usually offerWhy it matters commercially
Saudi Arabia, UAE, Qatar, Kuwait, Bahrain, OmanGulf and Najdi Arabic, with English mixed inModern Standard Arabic voices, read-style; dialect and code-switching weakHigh-income markets with phone-heavy banking, government and hospitality
SwitzerlandSwiss German dialectsStandard German with a Swiss accent at best; the dialects have no written standard and no stock voicesBanking, insurance and healthcare that speak dialect to customers
LuxembourgLuxembourgishAbsent or nearly soA financial centre where the national language has almost no speech tooling
NorwayNynorsk and regional dialectsBokmål voices exist; Nynorsk and dialects thin or absentPublic sector and regional services that must use them
Iceland, Faroe Islands, GreenlandIcelandic, Faroese, GreenlandicIcelandic has a few voices after public investment; Faroese and Greenlandic largely absentSmall, wealthy, and obliged to serve citizens in the language
Malta, Ireland, Wales, Basque CountryMaltese, Irish, Welsh, BasqueA voice or two, usually read-style, mostly from public programmesOfficial-language obligations in prosperous places
SingaporeSingapore English and code-switched Malay, Mandarin, TamilGeneric English voices; the local variety and the switching are not thereOne of the richest markets on earth, phone-heavy, multilingual by default
Hong Kong and TaiwanHong Kong Cantonese with English, Taiwanese HokkienCantonese voices exist; the mixing is weak; Hokkien largely absentFinance, retail and public services that speak it
IsraelHebrewVoices exist; choice and conversational quality lag the big languagesTechnology and finance with a demanding domestic market
Austria, Belgium, CanadaAustrian German, Flemish, Québec FrenchVoices exist but few; national accents flattened toward the larger standardCustomers notice when a brand sounds foreign

If your market is on this list, the stock voice will be the thing your customers remember, and not fondly.

What custom training changes

A custom voice is built from your own recordings: a speaker you choose, recorded with consent and the rights settled, in the variety of the language your customers actually speak, with a pronunciation dictionary for your names and products. The result is yours, runs where you need it, and does not change when a vendor updates a catalogue.

The same applies, with more force, to recognition. If the stock voice cannot speak Gulf Arabic, the stock recogniser cannot hear it either, and a voice agent needs both. Our deepest work is in Gulf and Najdi Arabic for exactly that reason. The method is the same for any of the rows above: collect, transcribe, train, measure on your own audio.

Before you commission anything

Three checks that take an afternoon:

  1. Listen to the current stock voices in your language on a phone, not on headphones. Play them to three people who speak the variety. If they laugh, you have your answer.
  2. Count the English words in a hundred real customer sentences. If it is more than a handful per sentence, code-switching is your main requirement, and it is the one stock voices handle worst.
  3. Find your speaker. A custom voice is only as good as the person and the recording. Settle the consent and the rights before the first session, in writing.

Then bring us the recordings. Twenty minutes is enough to say what can be done with them, and what it would take.

Author

Zingaro AI

The team that builds and runs AI operations for clients in the United States and the Middle East.

About the teamRSS

Share
LinkedInX
Have a job like this?

Twenty minutes is enough to say whether we can take it.

Book a call
Read next
All pieces
Book a call

Bring us the job you keep postponing.

Twenty minutes is enough to say whether we can take it.

A pilot starts within 5 working days of agreed scope · Nothing upfront · No seat licences