Skip to content
TrackHeros

Guide

Why Arabic–English meetings break most transcription, and what a two-pass transcript does about it

An explanation for anyone who runs client meetings in Arabic and English on Google Meet or Microsoft Teams: how code-switching works in Levantine and Gulf business talk, why single-language recognition mangles the technical words, why live captions are a scaffold rather than a record, what to check in a bilingual transcript, and dialect versus Modern Standard Arabic. Then how the TrackHeros two-pass transcript works.

Updated 2 September 2026 · 6 min read

Arabic–English business meetings break most transcription because speakers switch language inside a sentence, and a recogniser built for one language has to force the other language's sounds into its own vocabulary. A two-pass transcript avoids that by using the platform's live captions only for speaker names and timing, then transcribing the recorded audio with both languages expected, and treating that second pass as the record. This guide is for anyone who runs client meetings in Arabic and English on Google Meet or Microsoft Teams and needs the transcript to be right.

How does code-switching actually work in a business meeting?

In a Levantine or Gulf working meeting, nobody speaks Arabic for a paragraph and then English for a paragraph. The switch happens inside the sentence, and it follows a pattern. The frame of the sentence is Arabic. The technical noun is English. The verb that acts on the noun comes back in Arabic.

خلينا نعمل deploy عالـ staging بكرا الصبح، بس لازم نعمل merge للـ branch أول.

Roughly: "Let's deploy to staging tomorrow morning, but we have to merge the branch first." Every technical word is English, every word that holds the sentence together is Arabic, and the English nouns take Arabic articles and prepositions as if they had always been there: 'al-staging, lil-branch.

Numbers, dates and money do the same thing in the other direction. A developer will say a deadline in English and a price in Arabic, or read a ticket number in English digits and then explain it in Arabic. Names of people, companies and systems are pronounced however the speaker first heard them. None of this is carelessness. It is the ordinary register of technical work in the region, and a transcript that cannot follow it is not a transcript of the meeting.

Why does single-language recognition fail here?

A speech recogniser does two things at once: it turns sound into candidate words, and it chooses between the candidates using what it expects the language to be. That second part is where code-switching breaks it.

Set to Arabic, the recogniser hears "deploy" and has to produce an Arabic word, so it produces the nearest Arabic-sounding one, with full confidence. Set to English, it hears khallina na'mel and returns three English words that were never said. In both directions the technical terms, the ones your action items depend on, are the words that get mangled, because they are the ones in the other language.

Language detection does not rescue this. Most recognisers detect a language per segment or per utterance, then decode that segment in one language. A switch inside a sentence is shorter than the segment, so the minority language loses. The result reads fluently and is wrong, which is worse than gibberish, because nobody checks fluent text.

Why are live captions a scaffold, not a record?

Live captions are produced as the meeting happens, from the audio stream, with no chance to look ahead and with one expected language at a time. They are good at exactly two things: who is speaking, because the platform knows which participant's microphone is live, and when, because each line is stamped as it arrives.

They are bad at everything code-switching needs. They cannot revisit a sentence once the second half has changed its meaning, they mishandle names, and they drop the fast, quiet aside where the real decision was made. Use captions for structure: the speaker names, the timing, the shape of the conversation. Do not send them to a client as the transcript.

What does a two-pass transcript do?

It separates the two jobs.

Pass one: captions Pass two: audio
Source The platform's live captions The full recording
Runs During the meeting After the meeting
Expected languages One at a time Both, together
Good for Speaker names, timestamps, structure The words themselves, including mid-sentence switches
Its role Scaffold The record

The second pass transcribes the whole recording knowing that Arabic and English will both appear, so "deploy" inside an Arabic sentence is a word it expects rather than noise it has to explain away. It can use the whole sentence, and the sentences around it, to decide what was said. The caption timeline is then used to put the right name on each line, so the result has the captions' speakers and the audio pass's words.

It costs more, because the recording is processed twice, and it takes longer than captions. For an English-only meeting the first pass is often enough; for a bilingual one the second pass is the transcript.

What should you check in a bilingual transcript?

Even a good transcript needs a five-minute read before it goes anywhere. Check, in this order:

  1. Technical nouns. Are the system names, ticket ids and product terms spelled the way the team spells them? A transcript that writes "UBVA" for a client called "UBA" will do it in every meeting until you tell it otherwise.
  2. Numbers, dates and money. Read every one against the recording. These are the lines that end up in a proposal or an invoice.
  3. Names. People, companies, vendors. Names are the words no language model expects.
  4. Attribution on the action items. Who was actually speaking when the commitment was made. A line labelled "Unknown" is a line to check, not a line to guess.
  5. The switch points. Find the sentences where the language changed and read those against the audio. That is where a single-language pass would have failed, and where a two-pass transcript earns its keep.
  6. Consistency of the output language. If the notes are written in English, was the Arabic rendered the same way throughout, or translated in one place and transliterated in another?

Keep a list of the project's vocabulary as you go. It is the cheapest improvement any transcription gets.

Dialect or Modern Standard Arabic?

Modern Standard Arabic is the register of news, contracts and school. Nobody runs a stand-up in it. Meetings happen in Levantine, Gulf, Egyptian or another dialect, and the dialects differ from the standard in the words that hold a sentence together: negation, question words, the way the future is formed, and the pronunciation of letters like qaf and jim.

A system that mostly knows the standard will either "correct" the dialect into words that were not said or fail on the small words entirely. When you check a transcript, the small words are the tell. If "we won't" and "we will" are being confused, or a question has become a statement, the dialect is not being heard. The technical nouns will be fine; it is the Arabic in between that carries the difference between "yes, next week" and "no, not next week".

How TrackHeros does it

TrackHeros makes transcripts in two passes. The platform captions come first, for the real speaker names. Then an audio pass over the recording produces the authoritative transcript, and that pass is what handles Arabic and English switching mid-sentence. Speaker attribution is never guessed: a line that cannot be attributed is labelled "Unknown" rather than assigned to the most likely person.

Languages are set per client: the languages expected in the meeting (English and Arabic by default) and the language the transcript and notes are written in (English by default). Each transcript records which languages were detected and flags mid-sentence switching, and recurring series remember it. An English-only weekly call skips the audio pass after it has been seen clean several times; one sighting of a second language turns it back on permanently. Correction rules per client fix vocabulary ("UBVA" becomes "UBA") and instructions ("ticket ids look like PROJ-123") for future meetings, and a past meeting is fixed by re-summarising. More languages are not yet available. The feature is described at /features/arabic-english/, the notes it feeds at /features/meeting-minutes/, and the developer's day-to-day at /for/freelance-developers/.

Not yet available

  • More languagesNot yet available

Questions

Frequently asked

What is code-switching?

Switching between two languages inside one conversation, often inside one sentence. In Arabic–English business meetings the frame of the sentence is usually Arabic and the technical nouns are English, with Arabic articles and prepositions wrapped around them. It is a normal register, not a mistake.

Can live captions handle Arabic and English in the same sentence?

Not reliably. Live captions decode the audio stream as it arrives, with one expected language at a time and no chance to revisit a sentence. They are good at who spoke and when, which is why they make a useful scaffold, and poor at the words around the point of the switch.

Which language are the transcript and the notes written in?

In TrackHeros, the language you choose per client, English by default. The languages expected in the meeting are a separate per-client setting, English and Arabic by default. Only English and Arabic are promised today; more languages are not yet available.

Does it matter whether the meeting is in a dialect or Modern Standard Arabic?

Yes. Meetings happen in Levantine, Gulf, Egyptian or another dialect, and the dialects differ from the standard in the small words: negation, question words, how the future is formed. When you check a transcript, those small words are where a system that mainly knows the standard goes wrong.

Related

Early access

Get in early.

We’re onboarding a small group at a time so we can get it right. Join the list and we’ll reach out.

Get early accessPrivate beta · Google Meet and Microsoft Teams