ai translation14 min read

40+ languages translated live, then deleted at hang-up: multilingual transcription is the layer that keeps the meeting

DeepL launched Voice-to-Voice on April 16, 2026 across 40+ languages, and the same product's FAQ says meeting data is 'processed temporarily in memory and deleted once the call ends.' On August 28 the Open ASR Leaderboard added its first Global South language and showed two models tied at 4.9 WER differing almost fourfold in how much their accuracy depends on where the speaker is from. Real-time meeting translation is becoming a commodity. The record is not.

K
Ken Jo
#multilingual-transcription#real-time-meeting-translation#meeting-translator#deepl#live-translation#meeting-notes#asr#multilingual-meetings

On April 16, 2026, DeepL launched Voice-to-Voice from Cologne: a real-time spoken translation suite covering virtual meetings, mobile and web conversations, group settings for frontline workers, and an enterprise API. It handles more than 40 languages, including all 24 official EU languages alongside Vietnamese, Thai, Arabic, Norwegian, Hebrew, Bengali, and Tagalog. In blind evaluations run by Slator and commissioned by DeepL, 96% of linguists preferred it to the native translation built into Google, Microsoft, and Zoom; DeepL Voice for Zoom scored 96.4 out of 100 against 87–89 for the competing platforms.

That is a serious product, and the launch is not the interesting part. The interesting part is a sentence in the FAQ of the same product page, which we read again on August 31, 2026:

"DeepL does not permanently store transcription or translation data. Meeting data is processed temporarily in memory and deleted once the call ends."

Read those two facts next to each other. The translation is excellent, and the translation is gone by dinner.

TL;DR

  • Real-time meeting translation crossed from hard problem to shipped commodity in 2026: 40+ languages from DeepL, 58 monolingual languages on Deepgram's Nova-3, live translation inside Google Meet.
  • Ephemerality is usually a deliberate privacy feature, not a bug — but it means the meeting itself leaves no reviewable artifact.
  • Multilingual transcription is the separate layer that keeps the source-language transcript, the translation, the speaker labels, and the summary in one reviewable record. This post covers why the source language specifically must survive, and how to run the meeting so it does.

We made a Google-shaped version of this argument in June, when Meet and Translate pushed live speech translation to consumers — see why multilingual meetings still need reviewable notes. Since then the market has answered a question we were only asking. Live translation works now. So the interesting question moved: what do you have on Monday?

Real-time meeting translation stopped being the hard part

For about a decade, live cross-language speech was the demo that never survived contact with a real meeting. Latency ate the turn-taking, accents broke the recognizer, and anything technical came out mangled. That era ended over roughly ten months.

Deepgram's Nova-3 listed 31 languages in a December 10, 2025 release note and 58 in its documentation by August 2026 — 39 documented additions in under nine months, roughly one new language every seven days. We walked through that expansion, and the fine print underneath it, in what "supported" actually hides in speech-to-text language coverage. Google pushed live speech translation into Meet. DeepL shipped a four-surface voice suite in April.

The quality claims are now specific enough to argue with, which is itself a sign of maturity. DeepL's commissioned Slator assessment reports a 4% error rate for DeepL Voice against a 17% average for competing meeting platforms, and 96.4/100 for Zoom and 96.3/100 for Teams. Treat vendor-commissioned numbers as vendor-commissioned. But the direction is not in dispute, and the honest reporting around the launch agrees: The Next Web, covering the April 17, 2026 announcement, described a live demo in Seoul running at a one-to-two sentence delay, with DeepL's chief product officer acknowledging that word-order differences between languages remain a fundamental constraint on how fast speech-to-speech can ever be.

A one-to-two sentence delay is good enough. It is not good enough to replace a professional interpreter in a treaty negotiation, and nobody claims it is. It is comfortably good enough for a Tuesday product sync between Warsaw, Seoul, and São Paulo, which is the meeting most of us actually have.

Here's the point: when a capability becomes this good this fast, the bottleneck moves. It has moved. It is now sitting one layer down, in what the meeting leaves behind.

By design, the translation does not outlive the call

Go back to that FAQ line, because it is not sloppiness. It is a security posture, and a good one. DeepL's own wording continues: all data is encrypted in transit, never used to train models, and "only persists on the local device of meeting participants." For a bank or a hospital procuring a meeting tool, "we keep nothing" is the answer that clears legal review fastest.

The same posture is common across the live-caption category, because storing multilingual meeting audio is exactly the kind of liability nobody wants to inherit. So the default across the industry is now: superb comprehension in the moment, and no artifact afterward.

Diagram comparing what remains after a multilingual meeting when only real-time translation is used versus when multilingual transcription is kept

During the call, the two approaches are indistinguishable. The difference is entirely on the other side of hang-up. Left column follows DeepL Voice for Online Meetings' own FAQ, retrieved August 31, 2026.

The failure mode this produces has a shape, and it needs a name. Call it decision drift: everyone understood the meeting perfectly, and three weeks later nobody can prove what was agreed. It is not a misunderstanding — the live translation did its job. It is that the only surviving copy of a decision made in two languages lives in five people's memories, in five different phrasings, in at least two languages, and memory reconstructs rather than replays.

Decision drift is expensive in a way that never shows up on the tool's invoice. It shows up as a rebuilt feature, a renegotiated deadline, a compliance question nobody can answer, an onboarding hire who has no way to learn what the team decided last quarter.

August 28: the benchmark that explains why you need the source language

If transcription were a solved, uniform commodity, keeping the record would be a storage problem and this post would be short. It is not, and the clearest recent evidence landed three days ago.

On August 28, 2026, Hugging Face and Voice Arena added the first Global South language to the Open ASR Leaderboard: Hindi, spoken by more than half a billion people, joining a multilingual tab that until then covered only European languages. They contributed four evaluation splits — Monsoon en-IN and hi-IN, public and private — comprising 4,888 speaker-disjoint speakers, with 12 recorded attributes per speaker, drawn from 428 districts and hundreds of distinct handset models rather than a studio.

Then they ran the leaderboard's own models against it, and the result is the most useful thing published about speech recognition this month.

Bar chart showing 0.18 points of word error rate spread across eight leaderboard models, against 0.46 and 1.68 points of spread within a single model across five regions

Same unit on one axis: how many WER points separate the compared things. Source: Hugging Face and Voice Arena, August 28, 2026.

Eight models on the leaderboard land between 4.81 and 4.99 WER on the Indian English public split. That is 0.18 points from best to worst — inside what five hours of audio can even resolve. Ranked on the corpus, as the write-up puts it, they are the same model.

Group the speakers by region and they stop being the same model. Rolling each speaker's native district up to its zonal council, openai/whisper-large-v3-turbo varies by 0.46 points across the five zones. mistralai/Voxtral-Mini-3B-2507, fourteen hundredths of a point behind it on the corpus average, varies by 1.68 — 4.38 WER in the Central zone against 6.06 in the East. Two systems indistinguishable on the leaderboard differ almost fourfold in how much their accuracy depends on where the speaker grew up.

And it is not that one region is simply harder. ibm-granite/granite-speech-3.3-2b is worst in the North, microsoft/VibeVoice-ASR-HF in the South, Voxtral in the East. If some zone were intrinsically difficult, every model would rank the zones the same way. They do not, which points at the models rather than the audio.

The operational reading for anyone running multilingual meetings: your transcription quality is not a single number you can look up. It is a function of who is in the room. The colleague whose accent your vendor sampled least is the colleague whose sentences will be wrong, and neither the leaderboard nor the vendor page will warn you. The only defense is a transcript a human can read and fix — which requires that a transcript exist.

Hindi has ten valid spellings of the same phrase. A translation picks one.

The second finding in that release is subtler and cuts closer to the argument.

English orthographic variation is bounded: British against American spelling, punctuation, digits against words. A normalizer maps most of it to one form. Hindi is not bounded that way. Everyday speech is heavily code-mixed, English-origin words have no settled Devanagari spelling, and compounds are written joined or separated by preference. A single phrase can have ten or more valid written forms, with no canonical side to map them to.

So the Hindi splits ship a lattice — for each span of the transcript, the set of written forms accepted as correct — and are scored with Orthographically-Informed Word Error Rate (OIWER), introduced by AI4Bharat, rather than plain WER. When the same hypotheses were rescored against a flattened single-reference version, error rates rose for every system, unevenly, and pairs of systems reversed rank. Under a single reference a model is partly rewarded for reproducing the annotator's spelling choices; under the lattice it is scored only on recognition.

Sit with what that implies. Even at the level of "what was written down," there is often no single correct string — there is a set of admissible ones, and picking one throws information away.

Translation is that same collapse, one level up and far more lossy. Every translated line is a decision among readings that the source did not have to make.

Diagram of one meeting note holding three aligned layers for the same utterance: source-language text, translation, and summary

The source line is the only one of the three you cannot regenerate from the others.

"On va essayer de le faire d'ici fin septembre" becomes "We'll try to get it done by the end of September," and something real has changed: essayer carries a specific French hedging register that "try" flattens. Was that a commitment or an intention? The translated line cannot tell you. The source line can. Six weeks later, when the date slips and two teams remember the sentence differently, the source line is the entire argument.

This is the asymmetry at the center of the whole topic. Translation is one-way. You can re-translate from the source as many times as you like, with a better model next year. You can never recover the source from a translation.

What survives after the call ends: real-time translator vs meeting translator with notes

Here is the comparison stated plainly, across the three ways teams currently handle a multilingual meeting.

What you have the morning afterReal-time translation layerHuman interpreterMeeting translator with notes
Live comprehension in the roomYesYes, highest qualityYes
Audio of the meetingOnly if recorded separatelyOnly if recorded separatelyYes
Source-language transcriptNoNo, unless separately transcribedYes
Translation aligned to the sourceNoNoYes
Who said which lineNoNoYes, speaker labels
Summary and action itemsNoInterpreter's notes, if anyYes, generated from the transcript
Searchable three months laterNoNoYes
Correctable when a line is wrongNoNot after the factYes, edit the record
Typical cost per hourSoftware subscriptionUS$100–200+, two interpreters for long sessionsSoftware subscription

Ephemeral-column behavior follows DeepL Voice for Online Meetings' published FAQ, retrieved August 31, 2026; interpreter rates are the widely published market range for professional conference interpreting and vary by language pair and market. The third column describes the meeting-translator-with-notes category generally, not one vendor.

The table's shape is the argument. Column one and column three are identical in the only row a buyer usually evaluates — live comprehension — and differ in every row that matters after the meeting ends. Vendor comparisons are almost always run on row one.

A live translator is a window. A meeting translator with notes is a window and a ledger. Teams shop for windows because that is the part they experience, and then run their quarter off a ledger they never bought.

How to run a multilingual meeting so the record stays usable

Six steps, in the order they occur in a meeting. None of them are exotic; the failure is almost always that nobody decided.

  1. Set the source language explicitly. Do not lean on auto-detect. Auto-detection has to commit to a guess in the first seconds, on the least contextual audio in the whole session, and a wrong guess degrades every downstream stage at once. Explicit selection also lets the engine route you to a dedicated monolingual model, which is generally stronger than a multilingual one — Nova-3, for example, documents 58 languages one at a time but code-switching in exactly 10.
  2. Separate the languages people will speak from the languages people will read. These are different lists, and treating them as one is the most common configuration error in the category. DeepL's own help center splits them exactly this way: "spoken languages" are what the recognizer accepts, "translation languages" are what captions can be displayed in, and the second list is far longer than the first.
  3. Decide before the call where the record will live, and confirm it is not the meeting platform. If your translation vendor deletes meeting data at hang-up — which is the default, and a defensible one — then no configuration inside that tool will produce a record. The capture has to be a deliberate, separate choice.
  4. Keep speaker labels on. A translated line with no attribution is unusable for any decision it might later support. "We'll do it by September" is a fact only once you know who said it.
  5. Correct names and domain terms in the first five minutes, live. Proper nouns, product codenames, and acronyms are where transcription reliably fails, and they are also where a wrong line does the most damage to a summary generated later. Fixing them while the meeting runs costs seconds; fixing them in a document nobody re-reads costs nothing because nobody does it.
  6. Generate the summary from the source-language transcript where you can, not from the translation. Summarizing a translation compounds two lossy steps. This one is our recommendation rather than a benchmarked result, but the mechanism is straightforward: each stage discards information, and stacking them discards more.

Then keep the three artifacts in one record. A transcript in one tool, a translation in a chat log, and a summary in a doc are three files that will disagree with each other within a month, and the disagreement is the thing you were trying to prevent.

The bottom line: comprehension was the bandwidth problem, the record is the memory problem

Real-time meeting translation solves a bandwidth problem — moving meaning across the table fast enough that the conversation still feels like a conversation. In 2026, with 40+ languages at a one-to-two sentence delay and 96% linguist preference in a blind test, that problem is substantially solved and getting cheaper every quarter.

Multilingual transcription solves a memory problem, and nothing about the bandwidth solution touches it. Different artifact, different retention policy, different failure mode. A team that buys only the first and assumes it got the second will keep having excellent meetings it cannot account for.

So the question worth bringing to your next vendor call is not "how many languages does it translate." Every serious vendor now answers that with a number above 40. The question is: when the call ends, what still exists — and in whose language?


Where Telli.sh fits: everything above is about the layer under the live translation, and that is the layer we build. Telli.sh runs live transcription with translation into 44 target languages, with the interface itself available in 15, so a participant in Warsaw reads a Korean standup in Polish while it happens. The part that matters the morning after is that all three artifacts stay in one note: the source-language text of what was actually said, the translation aligned to it, and the summary generated on top — reviewable, correctable, and searchable long after the call ends.

Start a live translated meeting note

Sources


Back to Blog