Image for Training Slides Built by AI: Only 1.3% Get a Voice

Training Slides Built by AI: Only 1.3% Get a Voice

Only 11 of 860 AI-built training slide decks became a narrated module. Data on the languages training slides are written in, and why the voice never gets attached.

Enterprise training has quietly become a multilingual problem. The compliance module written in Toronto is delivered in Ho Chi Minh City, the safety induction drafted in London is read in São Paulo, and the product update built for a US sales team lands in a Riyadh office three weeks later. Voice AI vendors have spent two years promising to close that gap. The data suggests the gap is not where most of them are looking.

We examined 860 training decks built by 788 distinct users on the presentation platform ChatSlide between 9 June and 6 September 2026. Each deck carries a label for the kind of training it is and the language it was written in, which makes it possible to ask a question the vendor surveys in this space rarely answer: when an enterprise builds training material, what language is it in, and does anyone ever give it a voice?

Only 1.3% of training decks ever get narrated

Of 860 training decks built in that window, 11 were turned into a narrated video — 1.3%. The other 98.7% exist as slides, and slides only.

That number deserves to be read carefully, because it is easy to misread as a failure of demand. It is not. The people building these decks had narration available to them, one click away, in the same product. They built the training and stopped.

The reason is almost certainly workflow rather than appetite. A training deck is finished when it can be handed to whoever is delivering the session. Narration is an extra decision — which voice, which language, whether the phrasing that reads well on a slide also sounds right spoken aloud — and it arrives at exactly the moment the author considers the work done. Anything that lands after the perceived finish line gets skipped, however cheap it is.

For voice AI vendors this is the more useful finding. The barrier is not cost, quality, or awareness. It is that narration is positioned as a post-production step in a workflow that has no post-production stage.

The material is already multilingual

Bar chart of languages enterprise training decks are written in: English 290 (34%), British English 129 (15%), Vietnamese 86 (10%), Continental Spanish 39 (5%), Brazilian Portuguese 33 (4%), Spanish 33 (4%), Arabic 31 (4%), other or unspecified 219 (25%). n = 860 decks by 788 users, June to September 2026.
Bar chart of languages enterprise training decks are written in: English 290 (34%), British English 129 (15%), Vietnamese 86 (10%), Continental Spanish 39 (5%), Brazilian Portuguese 33 (4%), Spanish 33 (4%), Arabic 31 (4%), other or unspecified 219 (25%). n = 860 decks by 788 users, June to September 2026.

Language Decks Share
English 290 34%
British English 129 15%
Vietnamese 86 10%
Continental Spanish 39 5%
Brazilian Portuguese 33 4%
Spanish 33 4%
Arabic 31 4%
Other or unspecified 219 25%

n = 860 training decks by 788 distinct users, 9 June – 6 September 2026. Percentages are of all training-labelled decks and are rounded. Source: ChatSlide product data.

Two things stand out. English of some variety accounts for 49% — a plurality, not a majority. And Vietnamese, at 10%, sits third, ahead of every European language in the set.

That last figure is worth dwelling on. Vietnamese is not a language most enterprise voice stacks treat as a first-class citizen; it typically appears somewhere in a long tail of "additional languages" with thinner voice inventories and weaker prosody models. Here it is the third most common language enterprise training is written in. The mismatch between where training content actually originates and where voice coverage is deepest is a real commercial gap, and it points in an unfamiliar direction.

What kind of training this is

The distribution of training types explains why narration matters more here than in a generic deck.

Training type Decks Share
General 228 27%
Skills 189 22%
Onboarding 149 17%
Product 91 11%
Technical 84 10%
Compliance 44 5%
Leadership 39 5%
Safety 18 2%
Sales training 18 2%

Onboarding, compliance, and safety — 211 decks, roughly a quarter of the set — share a property the others do not: they are delivered repeatedly, to cohort after cohort, usually by someone who did not write them. That is the canonical case for synthetic narration. A deck delivered once by its author needs no voice. A safety induction delivered to every new hire for the next two years, in four languages, is a voice problem wearing a slide deck's clothes.

Compliance and safety together are only 7% of the sample, which is lower than the attention those categories get in enterprise voice AI marketing. The larger opportunity by volume is duller: onboarding and skills training, 338 decks between them, delivered to people who will never meet the author.

The delivery gap, stated plainly

Put the three findings together and the shape is clear. Enterprise training is being authored in many languages, a meaningful share of it is built to be delivered many times by someone other than its author, and almost none of it is given a voice.

The tooling response has generally been to sell narration as a separate product — record here, upload there, sync the timing yourself. The 1.3% figure is what that friction costs. Platforms that generate the deck and the narration in the same pass remove the handoff entirely; an AI training presentation that can attach a narrated presenter to the deck it just built is solving the workflow problem rather than the audio problem, which is the one the data says is binding.

This matters beyond convenience. In the United States, OSHA's position on training standards is that instruction must be presented "in a manner that employees receiving it are capable of understanding" — the agency has held that an employer cannot rely on a work rule it never communicated to a non-English-speaking employee. A slide deck in the wrong language is not a communicated rule. Neither, arguably, is one delivered by a manager reading unfamiliar material aloud from a script.

The standards side assumes the audio exists. The W3C Voice Browser Working Group has spent years on the premise that the web is something you can hear and speak to as well as see, and the multilingual compliance-auditing practice growing up around enterprise voice deployments assumes there is a recording to audit. For 98.7% of this sample, there is not.

Caveats worth stating

This is one platform's data and it describes the organisations that chose that platform. Users self-select the training label, the language field records the language the deck was authored in rather than every language it was later delivered in, and three months is a short window. The Vietnamese share in particular may reflect the platform's own regional distribution more than the global market.

None of that changes the ratio at the centre of it. Eleven of 860. Whatever the composition of the sample, the finished artefact of enterprise training today is a silent slide deck, and the voice layer everyone is selling has not yet been attached to it.

Author

Aisha Kamara

2026/09/07

Share this article