search
location
Request a Demo

AI medical scribes in Arabic: what actually matters

An AI medical scribe that works in English will not necessarily work in Arabic, and the reasons are structural rather than a matter of tuning. A Gulf consultation is rarely in one language, rarely in the written form of Arabic, and rarely free of English drug and procedure names. This page sets out what breaks, what to test before you buy, and what to ask for beyond a readable note.

By NANO Health Suite Clinical & Coding Team · Last updated

On this page

  1. Why Arabic is harder than English here
  2. What to test before you buy
  3. The dialect question, answered honestly
  4. The note is not the finish line
  5. Regulation, residency and integration in the Gulf
  6. How DoctorSense answers this

Why Arabic is harder than English here

Four properties of Arabic clinical speech each defeat a system built for English, and they compound rather than add.

Diglossia. The Arabic a clinician writes is not the Arabic a clinician speaks. A model trained on written Modern Standard Arabic — news, documents, books, which is where most Arabic text data comes from — meets a consultation conducted in regional speech and performs far worse than its published benchmarks suggest.

Code-switching. The sentence is Arabic; the drug is English. So is the procedure, frequently the anatomy, and almost always the abbreviation. A system that has to decide on one language per utterance will be wrong several times per consultation, and wrong in exactly the places that matter clinically.

Diacritics are omitted. Written Arabic normally leaves out the short-vowel marks, so words that are distinct in speech are identical on the page. Resolving them takes context, and in a clinical note the context is medical rather than general.

Direction and rendering. The note is right-to-left, the embedded English is left-to-right, and the numerals may be either. A note that is correct as data and mangled as a document is still a note nobody will sign.

What to test before you buy

A demonstration proves that a vendor can produce a good note from a recording they chose. It proves nothing about yours. Six tests, in the order that eliminates fastest:

1. Your own audio, unprepared. Record real consultations in your own rooms, with your own microphones and your own background noise, and have the vendor run those. A scribe evaluated only on clean studio audio is evaluated on the easiest problem it will ever face.

2. Code-switched sentences specifically. Not an Arabic consultation and an English one — a single sentence carrying an Arabic clinical frame and an English drug name. That is the normal case in the Gulf and the one most likely to break.

3. The speakers you actually have. Clinicians in the region come from many countries and speak accordingly, as do patients. Test with the range your clinic really sees rather than one representative speaker.

4. What happens when it is unsure. A system that guesses confidently is more dangerous than one that flags. Ask to see how uncertainty is surfaced to the clinician, because that behaviour is what determines whether review is real or a formality.

5. Structure, not prose. A paragraph is easy; a note organised into the sections your record expects, with the right content in each, is the actual product. Ask how many sections and check they are the ones your specialty uses.

6. Codeable output. See the next section — this is the test most buyers skip and most regret skipping.

The dialect question, answered honestly

Vendors advertise dialect coverage as a list of country names. Treat the list as a claim to be tested rather than a specification, for a simple reason: there is no agreed benchmark behind it. Nobody publishes what "supports the Gulf dialects" means, what error rate it corresponds to, or on what audio it was measured.

The question that produces a usable answer is narrower. Ask for word error rate on YOUR recordings, and separately ask for accuracy on the clinical entities — drugs, doses, diagnoses, negation — because those are where an error changes the meaning of the note rather than its readability. A scribe can score respectably overall and still lose the one word that mattered.

Negation deserves its own test. "No chest pain" and "chest pain" differ by one word, and in a structured note that word decides what is recorded as a finding. Ask specifically how negation is handled in Arabic, where it can be expressed in more than one construction.

Finally, ask what happens to accuracy over time. A model that is not being retrained on the speech it actually meets will drift away from your clinicians rather than towards them.

The note is not the finish line

A scribe that produces a beautiful note and nothing else has solved the smaller half of the problem. In a region moving to classification-based funding, the same encounter has to survive being coded, grouped and billed — and the note is the evidence for all three.

That is a documentation-integrity question, not a transcription one. The note has to carry the specificity a coder needs: the severity, the comorbidity, the laterality, the complication documented as a complication. A note that reads well to a clinician and vaguely to a coder produces a lower-weighted group and a smaller payment, and nobody notices because nothing failed.

So the question to ask a scribe vendor is not only "is the note accurate" but "is the note codeable". Ask which classifications the output is aligned to, and whether anything downstream checks the note against what the coder will need before the claim is built.

This is where an ambient scribe and clinical documentation integrity stop being separate purchases. The scribe writes it; documentation integrity defends it.

Regulation, residency and integration in the Gulf

Three practical constraints decide whether a technically good scribe can actually be deployed here.

Regulatory acceptance. Saudi Arabia regulates health data and AI actively, and a vendor that can point to recognition from the relevant national authority is answering a procurement question that will be asked regardless. Ask what the recognition covers and when it was granted rather than accepting the logo.

Data residency. Ask where the audio goes, where it is processed, where it is stored, for how long, and who can reach it. Ask it about the transcription and the audio separately, because the answers are often different.

Integration without migration. The scribe has to reach the record you already run. A tool that requires replacing the HIS to work is not a scribe deployment, it is a record migration with a scribe attached — and it will be scheduled accordingly. Ask how it attaches to the existing system and what format it writes back in.

How DoctorSense answers this

DoctorSense is NANO Health Suite’s ambient clinical documentation product, and these are the claims it makes on its own page rather than a summary written for this guide.

Language. Medical speech recognition in Arabic and English with full code-switching support — built for consultations that move between the two rather than for one language at a time.

Output. A structured clinical note organised into 33+ clinical sections, following ICD-10, SNOMED-CT and LOINC, exported as HL7, FHIR or JSON into Epic, Cerner, Meditech and local systems.

Recognition and scale. Accredited by SDAIA and in the hands of more than 30,000 healthcare professionals, and compliant with Saudi eHealth regulations.

Deployment. A one-click Chrome extension over the HIS or EHR already in use, with go-live inside 24 hours and no migration.

The coding loop. DoctorSense writes the note at the source and NANO AI CDI 360 validates, codes and strengthens it downstream — which is the answer to the "is the note codeable" test above, and the part a standalone scribe leaves to you.

Frequently asked questions

Does AI medical transcription work in Arabic?

Yes, and the useful question is how well on your consultations rather than in principle. Arabic adds three difficulties English does not have — diglossia, heavy code-switching with English clinical vocabulary, and omitted diacritics — so published accuracy figures earned on English or on written Arabic do not transfer. Systems built for bilingual clinical speech exist; systems retrofitted from English generally disappoint on exactly the sentences that matter.

What makes Arabic harder than English for an AI scribe?

Mainly that the Arabic a clinician speaks is not the Arabic most models were trained on. Written Modern Standard Arabic dominates the available text, while consultations happen in regional speech. Add English drug and procedure names inside Arabic sentences, short vowels that are not written, and right-to-left rendering with left-to-right insertions, and four separate problems arrive together.

Can an AI scribe handle Gulf dialects?

Some handle regional speech considerably better than others, but a published list of dialect names is not evidence. There is no shared benchmark behind such lists, so ask instead for performance on your own recordings, with a separate figure for clinical entities — drugs, doses, diagnoses and negation — because that is where an error changes meaning rather than readability.

What is code-switching and why does it break scribes?

Code-switching is moving between two languages inside one conversation, often inside one sentence. In Gulf consultations the clinical frame is commonly Arabic while the drug name, the procedure and the abbreviation stay English. A system that commits to a single language per utterance will misread precisely the clinical terms, which is the worst possible place to be wrong.

How should we evaluate an Arabic medical scribe?

On your own unprepared audio, recorded in your own rooms. Include deliberately code-switched sentences, the range of speakers you actually see, and cases where the system should be unsure. Check the structure of the note rather than the fluency of the prose, check how negation is handled, and check that the output is codeable and not merely readable.

Is an AI medical scribe compliant in Saudi Arabia?

That depends on the vendor, and it is a procurement question worth settling early. Ask what national recognition the product holds, what that recognition actually covers, and where audio and transcripts are processed and stored. DoctorSense states that it is accredited by SDAIA and compliant with Saudi eHealth regulations.

Does the note the scribe produces get coded correctly?

Not automatically, and this is the test most buyers skip. A note can be accurate and still lack the specificity a coder needs — the severity, the comorbidity, the laterality — which under classification-based funding produces a lower-weighted group and a smaller payment with nothing appearing to fail. Ask which classifications the output aligns to and what checks the note before a claim is built.

Does DoctorSense work in Arabic?

Yes — natively. DoctorSense understands medical terminology in Arabic and English with full code-switching support, built for the region’s real multilingual consultations, and is compliant with Saudi eHealth regulations. It is accredited by SDAIA and used by more than 30,000 healthcare professionals, producing structured notes across 33+ clinical sections.

Read next

Talk to the NANO team

Tell us what you are trying to solve and we will show you how it works on your own episodes.

Contact us