ai translation16 min read

Google will translate your Vietnamese meeting in real time. It will not write a word of it down.

Google Meet's translated captions cover 69 languages including Vietnamese, and Google Translate's Live translate covers 74. Meet's Transcripts feature and its Gemini note-taker each support exactly 8 languages, and Vietnamese is in neither. DeepL Voice can only transcribe spoken Vietnamese through a third-party provider that has to be enabled by an admin and cannot use auto-detect. A look at what live voice translation actually gives Vietnamese teams in 2026 — and what it never gives them.

K
Ken Jo
#vietnamese#live-translation#google-translate#google-meet#meeting-notes#sea#asr#multilingual-meetings

Type live voice translation into Google in Vietnamese and you land on a genuinely good answer. Google Translate's Live translate feature supports 74 languages, Vietnamese among them, on both Android and iOS. It runs with or without headphones, keeps working when you lock the screen, and offers a face-to-face mode that splits the phone in half so each speaker reads their own side.

It works. That is not the interesting part.

The interesting part is what happens on the same platform four hours later, when you want to know what was actually agreed. Google Meet's Transcripts feature — the one that produces a durable Google Doc — supports exactly eight languages, and Vietnamese is not one of them. Neither is it one of the eight supported by Meet's "Take notes for me," whose own help page adds that the feature "supports one language at a time" and that "multiple languages spoken in the same meeting aren't currently supported."

TL;DR

  • Vietnamese is supported by every Google surface that listens: Live translate (74 languages), Meet translated captions (69 languages). It is supported by neither Google surface that writes: Meet Transcripts and Meet's Gemini note-taker, 8 languages each.
  • DeepL Voice can display captions in Vietnamese but cannot transcribe spoken Vietnamese with its own models — Vietnamese sits in a third-party tier that an admin must enable and that cannot use auto-detect.
  • This post maps exactly where Vietnamese stops on each major platform as of August 31, 2026, and gives a six-step procedure for running a Vietnamese–English meeting that leaves a record behind.

We made the general-purpose version of this argument earlier today, looking at what happens when the translation is excellent and gone by dinner. This one is narrower and, for Vietnamese teams, worse: the gap is not just that live tools forget. It is that the remembering features specifically skip Vietnamese.

What "live voice translation" actually gives you in Vietnamese today

Start with the honest accounting, because Google's free tools are very good at the job they are built for.

Google Translate's Live translate ships four modes, per its Android help page retrieved on August 31, 2026: Listening (hold the phone to your ear or use headphones for real-time translation of someone else), Conversation (translations play out loud, speakers take turns), Text only (read, no audio), and Custom settings. There is a separate Face to face mode where the screen splits and each speaker sees both the transcription and the translation of their own language on their half. The microphone detects automatically when one language stops and the other starts, so nobody is tapping a button between sentences.

Two details deserve more credit than they usually get. First, the session survives multitasking: Google's documentation states that translation "continues seamlessly when you switch apps, minimize the application, or when the screen locks." Second, Vietnamese can be downloaded for offline use, and downloaded packs can be upgraded to a higher-quality version from the same screen — which matters in a country where a site visit or a factory floor is not a guaranteed-connectivity environment.

One limitation is worth stating plainly because it surprises people: this is a phone feature. Google's own help page for the web version says, in one sentence, "You can't talk and translate at the same time on Google Translate for the web." If your bilingual meeting happens on a laptop, Live translate is not in the room.

Here's the point: for a conversation — a taxi, a vendor visit, a hallway chat with a visiting client — this is close to a solved problem, it costs nothing, and we would recommend it without hesitation. The failure is not in the conversation. It is in the meeting.

Gemini 3.5 Live Translate takes Meet from five languages to more than seventy

The picture is also moving fast, and moving in Vietnamese speakers' favor.

On June 9, 2026, Google announced Gemini 3.5 Live Translate, a speech-to-speech translation model covering more than 70 languages. As Slator reported on June 16, 2026, it is rolling out through the Gemini Live API and Google AI Studio in public preview, entering private preview in Google Meet for selected Workspace customers, and shipping globally inside the Google Translate mobile app on Android and iOS.

The number that matters for Meet is the before-and-after. Speech translation in Meet had been limited to five supported languages, with translation only to and from English. Gemini 3.5 Live Translate expands that to more than 70 languages and over 2,000 language-pair combinations, with broader availability planned later in 2026. For a Vietnamese team, a five-language English-pivot system was effectively no system; 2,000 pairs is a different product.

Google is positioning the model narrowly, which is a good sign rather than a bad one. Per Slator, the documentation contrasts Live Translate with Gemini's broader Live Agent capabilities: it functions as a real-time translation pipeline, accepts audio-only input, and deliberately omits function calling, search grounding, tools, and system instructions. It is an interpreter, not an assistant — the same distinction OpenAI drew when describing GPT-Realtime-Translate. Google also added a "listening mode" on Android that streams translated audio through the earpiece, so you can follow a translation without visibly wearing headphones in a meeting.

So the live layer for Vietnamese is improving quickly and from several directions at once. Now the other question.

The record line: where Vietnamese stops

Every platform has a line. On one side are the features that help you understand speech as it happens. On the other are the features that turn speech into something you still have next week. Call it the record line — and for Vietnamese, it is exactly where support disappears.

Diagram showing Vietnamese supported across Google's listening features and absent from its writing features

Language counts are Google's own, from four help pages retrieved August 31, 2026. Vietnamese appears in both listening features and neither writing feature.

Here is the same information as a table, because the asymmetry is easier to argue with when it is lined up.

Google featureVietnamese supported?LanguagesLeaves an artifact after the call?
Google Translate — Live translateYes74No
Google Meet — translated captionsYes69On screen only; embedded in the clip if you record captions
Google Meet — TranscriptsNo8Yes, a Google Doc
Google Meet — "Take notes for me"No8, one at a timeYes, notes in Google Docs plus an emailed recap
Google Meet — Gemini 3.5 Live TranslateRolling out5 today, 70+ in private previewNot stated

All rows from Google Help pages and Slator's June 16, 2026 report, retrieved August 31, 2026.

The eight languages shared by Transcripts and "Take notes for me" are identical: English, French, German, Italian, Japanese, Korean, Portuguese, and Spanish. It is the same list twice, and it is a list of large European and East Asian markets. Vietnamese — roughly 97 million speakers, by Samsung's own count in its 2024 write-up of the work — is outside it, alongside Thai, Indonesian, Hindi, and Arabic.

There is one more constraint that hits bilingual teams specifically, and it applies even to teams working entirely in supported languages. Google's help page for "Take notes for me" states that the feature supports one language at a time and that multiple languages spoken in the same meeting are not currently supported. A Vietnam–US call that switches between Vietnamese and English is disqualified twice over: once for Vietnamese, and once for switching.

To be fair to Google: translated captions do have a persistence path. If you record the meeting and select Record captions, the captions are embedded into the clip. And you can scroll back through translated captions during the session. But burned-in captions on a video file are not a searchable record — you cannot grep a pixel, you cannot correct a name, and you cannot generate a summary from it. Translated captions in Meet are also restricted to paid Workspace editions: Business Standard, Business Plus, Enterprise Standard, Enterprise Plus, Enterprise Starter (available until June 30, 2025), and Google AI Pro for Education.

DeepL will show you Vietnamese captions. Transcribing spoken Vietnamese is a different tier.

DeepL Voice is the strongest Western alternative, and its Vietnamese story is the most instructive one in this whole post — because DeepL documents the distinction better than anyone.

DeepL splits language support into two lists. Spoken languages are what the recognizer accepts. Translation languages are what captions can be displayed in. Vietnamese appears in the translation list, which runs to 41 languages. Now look at the other side: DeepL's own models transcribe 18 spoken languages — Chinese (Mandarin), Czech, Dutch, English, French, German, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Romanian, Russian, Spanish, Swedish, Turkish, and Ukrainian. Vietnamese is not among them.

Vietnamese is available as a spoken language, but only in a second tier of 20 languages transcribed by a third-party provider, Speechmatics. Three conditions come attached, all documented on DeepL's help page retrieved August 31, 2026:

  1. A team admin must enable third-party processing in the organization's admin account before any of these languages work at all.
  2. Auto-detect does not work. In DeepL's words: "Third-party languages do not work with Auto-detect language — they need to be manually selected as a Spoken language."
  3. To let participants pick their own spoken language, the host must enable an access code when connecting the meeting to DeepL Voice.

None of that is unreasonable — DeepL notes that Speechmatics operates under a zero-retention agreement, so transcription data is deleted from their servers once processed. But stack the requirements and you get the practical result: a Vietnamese speaker joining a DeepL-equipped meeting is one un-ticked admin setting away from not being transcribed at all, and cannot rely on the system noticing them automatically. Meanwhile the retention posture is the same one we wrote about this morning — DeepL's FAQ says meeting data is "processed temporarily in memory and deleted once the call ends."

Samsung is the counterexample worth naming. Vietnamese was among the first 16 languages Galaxy AI shipped, and Samsung R&D Institute Vietnam published an unusually candid account on May 23, 2024 of why it was hard. Vietnamese has six distinct tones; the article's example is that ma, mả, and — ghost, grave, and mother — differ only by tone. Bui Ngoc Tung, the institute's ASR lead, described a model differentiating audio frames of "around 20 milliseconds" to decide which word a run of frames belongs to. When a Korean handset maker invests in Vietnamese tone modeling and a European translation company routes it to a subcontractor, that tells you something about how the market prices this language.

On the speech-recognition supply side, the picture is better than the note-taking side: Deepgram's Nova-3 lists Vietnamese as vi in its models and languages documentation, retrieved August 31, 2026 — one of the 58 languages we counted in our analysis of what "supported" hides in speech-to-text coverage.

No vendor publishes a Vietnamese meeting-speech error rate, and the corpora explain why

We went looking for a number to put on Vietnamese transcription quality in meetings. There isn't one, and the reason is more useful than the number would have been.

Bar chart of labeled-hour counts for three Vietnamese speech corpora, annotated by speaking register

Labeled hours are not the whole story — register is. FLEURS is read speech; PhoWhisper's training set spans accents; VietSuperSpeech is the first corpus built specifically for casual conversation. Sources dated in the list below.

The benchmark everyone quotes for Vietnamese is FLEURS, which provides roughly 12 hours of read-aloud speech per language. Read speech is not meeting speech. PhoWhisper, presented in the ICLR 2024 Tiny Papers track, fine-tuned Whisper on an 844-hour Vietnamese dataset chosen for accent diversity, which is a real step toward realism. VietASR, published May 23, 2025, took a different route — pre-training on 70,000 hours of unlabeled audio and fine-tuning on just 50 labeled hours — and reports outperforming Whisper Large-v3 and commercial systems on real-world data.

Then there is the paper that says the quiet part out loud. VietSuperSpeech, released in 2026, assembles 52,023 audio–text pairs totaling 267.39 hours of casual conversational Vietnamese, split into 240.67 training hours and 26.72 evaluation hours, comprising 13.8 million fully diacritically marked characters. Its stated motivation is the sentence to quote:

"While corpora such as VLSP2020, VIET_BUD500, VietSpeech, FLEURS, VietMed, Sub-GigaSpeech2-Vi, viVoice, and Sub-PhoAudioBook provide broad coverage of formal and read speech, none specifically targets the casual, spontaneous register indispensable for conversational AI applications."

Eight named Vietnamese corpora, and until this one, not one of them built for the way people actually talk in a meeting — interrupting, hedging, code-switching into English for the product name, trailing off. Whatever accuracy figure a vendor quotes you for Vietnamese was almost certainly measured on someone reading a script.

The operational consequence is simple and it is the same one we keep arriving at. If you cannot look up how accurate the transcription of your meeting will be, the only defense is a transcript a human can read and correct. Which requires that a transcript exist in the first place.

How to run a Vietnamese–English meeting so the record survives

Six steps, in the order they happen. Every one of them addresses a specific failure documented above.

  1. Decide which tool is doing comprehension and which tool is doing the record — and accept that they are different tools. This is the whole post in one line. Google Translate or Meet captions can handle the first job for Vietnamese today. Neither handles the second.
  2. Set Vietnamese as the spoken language explicitly; do not rely on auto-detect. On DeepL this is mandatory — third-party languages, Vietnamese included, cannot be auto-detected. Everywhere else it is still the better choice, because auto-detection commits to a guess in the first few seconds, on the least contextual audio of the entire session.
  3. Check the record feature's language list, not the platform's. The trap this post exists to describe is assuming that a platform which translates Vietnamese also transcribes it. In Google Meet, 69 versus 8 is the gap between those two assumptions.
  4. Name the code-switching problem before the call. If your team drops English product names, ticket IDs, and acronyms into Vietnamese sentences — and every offshore engineering team does — verify that your record tool tolerates two languages in one session. Meet's note-taker explicitly does not.
  5. Fix proper nouns in the first five minutes, live. Vietnamese diacritics and English product codenames are the two places transcription reliably breaks, and they are also the two places a wrong line does the most damage to any summary built on top of it. Correcting during the meeting costs seconds.
  6. Keep the Vietnamese source text, the English translation, and the summary in one record. Not three files in three tools. Translation is one-way: you can re-translate from the Vietnamese next year with a better model, but you can never recover the Vietnamese from the English.

If what you want is the step-by-step product walkthrough rather than the market analysis, our June guide for Vietnamese teams covers turning a bilingual meeting recording into a reviewable AI note end to end.

The bottom line: Vietnamese has crossed the comprehension gap and not the record gap

For most of the last decade, the complaint about Vietnamese in international meetings was that the technology did not understand it well enough. In 2026 that complaint has largely expired. Between 74 languages in Live translate, 69 in Meet's translated captions, more than 2,000 language pairs coming to Meet via Gemini 3.5 Live Translate, Vietnamese in Nova-3, and Samsung shipping on-device Vietnamese interpretation since 2024, comprehension is handled.

What has not moved is the record. The two Google features that produce a durable artifact support eight languages between them, the same eight, and Vietnamese is in neither. DeepL will caption into Vietnamese but routes transcription of it to a subcontractor behind an admin toggle. The market has decided that Vietnamese is a language worth hearing and not yet a language worth writing down.

So the question to ask a vendor is not whether they support Vietnamese. Almost all of them now say yes, and almost all of them mean the listening half. The question is: which of your features support Vietnamese after the call ends — and can you show me the language list for that specific feature?


Where Telli.sh fits: the record layer is the layer we build. Telli.sh transcribes the meeting in the language it was actually spoken in, translates into 44 target languages including Vietnamese, and keeps all three artifacts in the same note — the Vietnamese source text, the translation aligned to it, and the summary generated on top. The interface itself ships in 15 languages, Vietnamese included, so a team in Da Nang and a client in Sydney read the same meeting in their own language and still share one correctable record.

Start a live translated meeting note

Sources


Back to Blog