Research

M3, measured on real meetings.

The Arabic meeting transcription system inside MeetriX reached a 1.62% character error rate on real multi-dialect Arabic business meetings with human-verified references, the lowest of the seven systems compared. Here is what that means for your meeting record, and everything we measured to say it.

Research

  • Published
  • VersionEvaluation Edition, September 2026
  • AuthorshipLisan Research
1.62%Character error rateMeasured on the same human-verified meetings as every compared system
28% to 83%Lower CER than each compared systemMeasured, approximate, against six speech APIs
about 45 sTo transcribe one hour of audioTeam-reported, about 80x real time on an L4 GPU, batch not live
7Systems scored on the same audio, reference and rulesM3 and six commercial speech APIs

What 1.62% means for a meeting record

A meeting record earns trust one detail at a time: the number in the budget line, the abbreviation in the technical update, the name of the vendor, the one-word answer that closed a decision. The evaluation behind this page scored exactly those details. It counted every character that a transcription system got wrong, dropped or invented against what a human verified was said, across real business meetings in several Arabic dialects with English terms inside the Arabic. Measured. M3 recorded 1.62%, which means roughly 16 wrong characters in every 1,000. The closest system, Google chirp_3, recorded 2.26%.

For a government committee, a board or a project team, the practical reading is in the examples further down: which systems kept 200 files and the scores 98 and 94, which kept the Sprint Planning and VAT terms in a form a search box will find later, which wrote SDK as SDK and which spelled it out in Arabic letters, and which returned nothing for a greeting shorter than a second. Those are the failures that make a rapporteur go back to the audio, and the evaluation was built to catch them.

M3 records the lowest character error rate of seven systems on the same human-verified meetings

Lower is better. Character error rate (CER) after defined lexical normalisation.

Same source audio, same human-verified reference, same scoring rules for every system. Every system received the same participant-separated audio, so speaker attribution is not part of the comparison. Evaluation snapshot September 2026. See "How it was measured" for the scoring rules.

Seven systems, one evaluation set, one set of rules. M3 at 1.62% is followed by Google chirp_3 at 2.26%, Cohere at 4.93%, ElevenLabs Scribe at 5.01%, Deepgram Nova-3 at 6.09%, Google latest_long at 9.01% and Sonix at 9.8%. The table adds character accuracy (100 minus CER) and the approximate relative reduction of M3 versus each system, from about 28% against chirp_3 to about 83% against Sonix. Read the reduction column as the gap on this evaluation set, not as a ranking of vendors by quality.

Character error rate and character accuracy for the seven systems, with M3's approximate relative reduction versus each. Lower CER is better. Measured on the same human-verified meetings; scoring rules under How it was measured. best in column
RankSystemCER (%), lower is betterCharacter accuracy (%)M3 vs this system
1M31.62 (best in column)98.38 (best in column)baseline
2Google chirp_32.2697.74about 28% lower
3Cohere4.9395.07about 67% lower
4ElevenLabs Scribe5.0194.99about 68% lower
5Deepgram Nova-36.0993.91about 73% lower
6Google latest_long9.0190.99about 82% lower
7Sonix9.890.2about 83% lower (best in column)

Where a record breaks, segment by segment

The cards show the human-verified reference and every system's output for individual segments, copied verbatim from the evaluation. Speaker labels are fictional placeholders. Each card carries the source clip: an excerpt of an internal Lisan meeting, the same audio every system transcribed. Each card was chosen because a reader without Arabic can still see the failure: an empty cell, a missing digit, a Latin term that survived or did not, a switch into another script. Where M3 itself makes a mistake, the caption says so. Under the scoring rules a phonetic spelling of an English term is not an error, so some of the differences affect readability and search rather than the CER.

Example 1English terminology
  • Speaker Kareem Saleh
  • Segment 38faf0f4_6
  • Clip length 42.66 s

SDK, rendering, editor, and Chrome extension

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

صباح الخير بالنسبة لي امبارح كان معظم الشغل مع تيم العربي الجديد بخصوص الـ SDK (preserved term) اتسكر تقريبا مشكلتين و إلهن علاقة بالـ rendering (preserved term) الـ highlights (preserved term) بالإديتور اللي معمل له الـ Configuration (preserved term) من عندن بصفى مشكلة وحدة طلبنا منهم اجتماع مشان نكمل معهم اليوم ان شاء الله، كمان امبارح صرت اشتغل على موضوع الـ Chrome extension (preserved term) بين كوركتو و Lisan (preserved term) ما كثير يعني خلصت هذا الموضوع، اليوم ان شاء الله بدي كمل مثل ما قلت بخصوص العربي الجديد وكمان الـ Chrome extension مع كوركتو، يعطيكم العافية.

Reference

صباح الخير، بالنسبة لي امبارح كان معظم الشغل مع تيم العربي الجديد بخصوص الـ SDK اتسكر تقريبا مشكلتين ولهم علاقة بالـ rendering للـ highlights بالـ editor اللي معمل له configuration من عندن بصفى مشكلة وحدة طلبنا منن اجتماع مشان نكمل معن اليوم ان شاء الله كمان امبارح صرت اشتغل على موضوع الـ Chrome extension بين كوركتا و Lisan ما كتير خلصت هذا الموضوع اليوم إن شاء الله بدي كمل متل ما قلت بخصوص العربي الجديد وكمان الـ (Chrome extension) مع العربي الجديد

Why this example matters A daily stand-up full of developer vocabulary. M3 keeps SDK, rendering, highlights, configuration, Chrome extension and Lisan in Latin script, which is what a search across your archive will find. ElevenLabs Scribe keeps them too, with heavy filler. chirp_3 is complete but phonetic, which the scoring rules accept.

  • Preserved: SDK, rendering, highlights, Configuration, Chrome extension, Lisan
  • ElevenLabs Scribe: unsupported text
  • Deepgram Nova-3: content dropped
  • Google latest_long: content dropped
  • Sonix: content dropped
  • Cohere: content dropped
Example 2Filler and punctuation
  • Speaker Rami Haddad
  • Segment a3ef0ea8_735
  • Clip length 20.81 s

Sprint Planning and VAT

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

صباح الخير يعطيكم العافية. آآآه (filler-like insertion) امبارح كان في عندنا متابعة باجتماعات الـ Sprint Planning (preserved term) آآه (filler-like insertion) كان في كمان متابعة مع آآه (filler-like insertion) الشي المطلوب من هيئة الزكاة والضرائب بالسعودية آآه (filler-like insertion) موضوع الـ آآه (filler-like insertion) الـ VAT (preserved term) ضريبة القيمة المضافة. كنت عم بتابع بهدول المواضيع يعطيكم العافية.

Reference

صباح الخير يعطيكم العافية امبارح كان في عنا متابعة باجتماعات السبرنت بلاننج كان في كمان متابعة مع الشي المطلوب من هيئة الزكاة والضرائب بالسعودية موضوع ال الزاد ضريبة القيمة المضافة كان كنت عم بتابع بهدول المواضيع يعطيكم العافية

Why this example matters A hesitant status update. M3 keeps Sprint Planning and VAT in Latin with the Arabic expansion of VAT, while Cohere prints five hesitation tags and several systems garble or drop the VAT term. Honest note: M3 transcribes the speaker's hesitation sounds here, where the reference has none, so this clip does not show M3 removing fillers.

  • Preserved: Sprint Planning, VAT
  • Cohere: unsupported text
  • ElevenLabs Scribe: unsupported text
  • Deepgram Nova-3: unsupported text
  • Google latest_long: unsupported text
  • Sonix: content dropped
  • Google latest_long: content dropped
Example 3Short speech
  • Speaker Omar Khalil
  • Segment ce9f3a3a_10
  • Clip length 0.71 s

A sub-second spoken greeting

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

كيفكم؟

Reference

كيفكم؟

Why this example matters A greeting shorter than a second. M3, Cohere, Sonix and chirp_3 match the reference exactly. Deepgram Nova-3 and Google latest_long returned nothing, which scores as deletions, and ElevenLabs Scribe produced an English word that was not spoken.

  • ElevenLabs Scribe: unsupported text
  • Deepgram Nova-3: content dropped
  • Google latest_long: content dropped
Example 4Numbers and terms
  • Speaker Tareq Faris
  • Segment 38faf0f4_8
  • Clip length 39.23 s

Files, quality figures, and model work

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

مساء الخير (not supported by the audio) امبارح خلصت الملفات كلهم طلعوا جوا الـ تبع الـ translated (preserved term) اكسل هلأ هنن حوالي 200 (preserved term) file (preserved term) الدقة طلعت يعني بالـ style (preserved term) حوالي الـ 98 (preserved term) يعني بالـ translation (preserved term) كجودة عم بـ 94 (preserved term) هلأ هدا بعد طبعا عدة تعديلات هلأ اليوم بدي ارفع هلأ هي الـ version (preserved term) وكان في كمان شغل على الـ NBOE (preserved term) يعني حاولنا نعمل صاين تيونينغ هلأ اليوم كمان حتابع

Reference

صباح الخير امبارح خلصت الملفات كلهم تبع لينا تبع TranslateX حوالي 200 file الدقه طلعت يعني بالـ style حوالي 98 يعني بالـ translation كجوده بال 94 هلا هذا بعد طبعا عده تعديلات هلا اليوم بدي ارفع هلا هي الـ version وكان في كمان شغل على MOE حاولنا نعمل fine-tunning، وسأتابع اليوم كمان

Why this example matters Three spoken figures: 200 files, a style score of 98 and a translation quality score of 94. M3, Cohere and chirp_3 keep all three. Elsewhere 94 became 40, 98 became eight, 200 was dropped, or 98 became 80 90. M3 is not clean here: it opens with the wrong greeting and misspells two English terms. One spoken first name is replaced by a placeholder in the transcripts.

  • Preserved: translated, file, style, translation, version, NBOE, 200, 98, 94
  • Cohere: unsupported text
  • ElevenLabs Scribe: unsupported text
  • Deepgram Nova-3: unsupported text
  • Sonix: content dropped
  • Google latest_long: content dropped
  • Google chirp_3: content dropped
Example 5English terminology
  • Speaker Kareem Saleh
  • Segment d974dd6c_4
  • Clip length 23.99 s

SDK, dark theme, web app, and monitoring

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

صباح الخير بالنسبة ليوم الخميس الشغل كان مع فضاءات كان في عندهم مشاكل العلاقة بالـ SDK (preserved term) والسججشنز اللي عم تطلع انحل اغلب الاخطاء صفيان وحدة بدنا نعمل كول عليها اليوم ان شاء الله كمان بدي اشتغل بنسخة الـ dark theme (preserved term) بالـ web app (preserved term) وكمان بدي اعمل مونتورنج للنسخة الجديدة تبع الاي اي ايجنتس يعطيكم العافية

Reference

صباح الخير، بالنسبة ليوم الخميس الشغل كان مع فضاءات، كان في عندهم مشاكل العلاقة بالـ SDK والـ suggestions اللي عم تطلع، انحل اغلب الاخطاء، صفيان وحدة بدنا نعمل كول عليها اليوم إن شاء الله، كمان بدي اشتغل بالـ بنسخة الـ dark theme بالـ web app، وكمان بدي أعمل monitoring للنسخة الجديدة تبع الـ AI agents، يعطيكم العافية.

Why this example matters Consecutive product terms. M3 keeps SDK, dark theme and web app in Latin and writes suggestions, monitoring and AI agents phonetically, which the scoring rules count as correct. ElevenLabs Scribe keeps every term in Latin. Sonix drops most of them, and Google latest_long garbles most of the rest.

  • Preserved: SDK, dark theme, web app
  • ElevenLabs Scribe: unsupported text
  • Sonix: content dropped
  • Google latest_long: content dropped
  • Deepgram Nova-3: content dropped
  • Google chirp_3: content dropped

How each system behaved

Qualitative. Beyond the CER, the reviewers noted how each system handled English terms inside Arabic, content at the start of segments, fillers and event tags, punctuation and repetition. No separate aggregate rates were calculated, and the matrix shows the reviewers' phrase in every cell so the strong, mixed, weak and not rated levels can be checked against the wording.

Google chirp_3 deserves the credit the review gives it: strong handling of English technical terminology within Arabic speech, low content loss, and substantially fewer empty outputs than the V1 configuration. On several example segments it is as complete as M3, with the difference limited to phonetic rather than Latin spellings. ElevenLabs Scribe preserved English terminology particularly well but produced frequent spurious filler-like vocalisations and one switch into an unexpected language or script. Cohere was highly accurate on Classical Arabic, while the review identified content omissions, inconsistent terminology, repeated punctuation and repetition loops. Deepgram Nova-3 showed participant-name errors and segment-start omissions. Google latest_long rejected many non-speech segments correctly but missed a small number of speech-bearing ones and frequently omitted English terms. Sonix showed errors in names and English terms, segment-start omissions and excessive sentence-final punctuation. All seven systems supported the Arabic dialects present in the evaluated meetings.

M3 recorded the lowest CER and, in the reviewers' notes, preserved business-critical content, numbers and terminology more reliably and did not show the same concentration of filler, tag or repeated-punctuation artifacts. Speaker identification was not evaluated for the external systems; each received participant-separated audio prepared by the M3 capture workflow and was scored on transcription only.

Nine criteria, seven systems, one phrase per cell. The M3 column carries the brand rule only; status colours never mark emphasis. The final row repeats the measured CER.

StrongMixedWeakNot rated
Observed output behaviour by criterion and system, from the reviewers' notes on the evaluation set. Levels follow one rule from the report's wording: strong where the reviewers state a positive result, weak where they state a defect, mixed where the note is qualified or says not separately scored or no recurring issue was isolated, and not rated where the criterion was not evaluated. The phrase is shown in every cell. Separate aggregate rates were not calculated.
M3CohereElevenLabs ScribeDeepgram Nova-3Google latest_longGoogle chirp_3Sonix
Classical ArabicStrongHighly accurate.StrongHighly accurate.StrongAccurate.StrongAccurate.WeakMany errors.StrongAccurate.StrongAccurate.
English inside ArabicStrongPreserved English technical and business terminology more consistently within Arabic speech.WeakEnglish technical and business terms were sometimes omitted, mistranscribed, or rendered inconsistently within Arabic speech.StrongEnglish technical terminology was preserved particularly well in the reviewed output.MixedEnglish-term performance was not scored separately; reviewed examples showed inconsistent rendering alongside segment-start omissions.WeakEnglish technical terms were frequently omitted.StrongEnglish technical terminology was preserved strongly within Arabic speech.WeakNames and English technical terminology were frequently transcribed incorrectly or omitted.
Text omissionStrongPreserves business-critical content more reliably and reduces the loss of important numbers, terminology, and organisation or product names.WeakMay omit or distort business-critical content, including numbers, technical terms, and organisation or product names.MixedNo separate omission rate was calculated; the principal qualitative issue was inserted filler-like content rather than a recurring segment-start omission pattern.WeakSubstantial content loss was observed, particularly at the beginning of source segments.MixedCorrectly rejected many non-speech segments, but missed a small number of speech-bearing segments and sometimes omitted English terms or numbers.StrongLow content loss was observed, with substantially fewer empty outputs than the V1 configuration.WeakContent loss was observed particularly at the beginning of source segments.
PunctuationStrongProduced clearer sentence boundaries without the same concentration of repeated-punctuation artifacts.WeakRepeated punctuation patterns that reduce readability were observed in the reviewed output.MixedNot separately scored; readability was affected more by filler-like insertions than by a recurring punctuation defect.MixedNot separately scored; no recurring punctuation defect was isolated in the qualitative review.MixedNot separately scored; no recurring punctuation defect was isolated in the qualitative review.MixedNot separately scored; no recurring punctuation defect was highlighted in the qualitative review.WeakPunctuation quality was weak, with excessive sentence-final marks observed in the reviewed output.
Spelling and terminologyStrongProduces more consistent business terminology for reading and search.WeakTechnical and business terms may appear inconsistently or in phonetic form.MixedEnglish technical terms were handled well, although filler-like insertions reduced overall transcript cleanliness.WeakParticipant names were frequently transcribed incorrectly.WeakEnglish technical terms were frequently missed, and some numbers were omitted.StrongEnglish technical terms were rendered reliably within Arabic speech.WeakPerformance was weak on participant names and English technical terminology.
Word repetitionStrongShowed substantially fewer repetition problems in the reviewed output.WeakToken and phrase repetition loops were observed in the reviewed output.MixedNo separate repetition rate was calculated; frequent filler-like vocalisations were the more prominent insertion issue.MixedNo recurring repetition issue was isolated in the qualitative review.MixedNo recurring repetition issue was isolated in the qualitative review.MixedNo recurring repetition issue was isolated in the qualitative review.MixedNo recurring word-repetition issue was isolated; punctuation was the more prominent readability problem.
DialectsStrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.
Speaker identificationStrongUses participant-linked audio channels and platform participant information. Platform-provided speaker attribution remained linked throughout the evaluated meetings.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.
Filler words and event tagsStrongDid not show the same concentration of filler-like, event-tag, or repeated-punctuation artifacts in the review.WeakFiller-like insertions, event-style tags, and repeated-punctuation artifacts were observed in the reviewed output.WeakFrequent spurious filler-like vocalisations such as “آآآ” were observed, and one segment switched to an unexpected language or script.MixedNo recurring filler-word or event-tag pattern was isolated in the qualitative review.MixedHandled many non-speech segments correctly, although a small number of speech-bearing segments were also rejected.MixedNo recurring filler-word or event-tag pattern was highlighted in the qualitative review.MixedNo recurring event-tag pattern was isolated; excessive punctuation was the more prominent artifact.
Character error rate (CER), lower is better1.62%4.93%5.01%6.09%9.01%2.26%9.8%

Speed, and what the figure covers

Team-reported. The M3 meeting transcription workflow processes a one-hour meeting in approximately 45 seconds on an L4 GPU, about 80x real time. This is batch throughput for recorded audio, measured on a meeting-duration basis. It is not live latency: MeetriX also supports live processing while a meeting is in progress, but live delivery is measured separately and no live latency figure is published here. For a minutes workflow the figure describes the transcription stage only: how long M3 takes to process an hour of recorded audio once it is available, not the time until a transcript is delivered, which depends on the rest of the pipeline. Speaker identity, timestamps and transcript order are preserved through processing (product capability).

One hour of meeting audio in about 45 seconds

About 80x real time, team-reported, L4 GPU.

Team-reported batch throughput for the M3 meeting transcription workflow on an L4 GPU. Describes processing of recorded audio, not live speech latency.

Inside MeetriX

Product capability. The benchmark scores transcription only. In MeetriX, M3 sits inside a longer workflow: audio is captured with participant identity attached from the moment of capture; silence, short turns and speech boundaries are handled before transcription; multi-dialect Arabic is transcribed with English names and terms written in their recognised business form; every segment stays linked to its speaker, meeting time and source audio; and a post-meeting stage produces the summary, key points, decisions and next steps that are emailed to the participants under your access, consent and retention policies. In the evaluated deployment, the transcription path called no external transcription service, and M3 can run within customer-controlled infrastructure so audio, transcripts and summaries stay inside your data-governance boundary. Platform-provided speaker attribution remained linked throughout every evaluated meeting in the M3 participant-channel workflow; this is a coverage statement, and no speaker-attributed error rate is reported.

Note

Product capability, not benchmark. Summaries, minutes and delivery were not scored in this evaluation. The CER figures on this page cover transcription only.

How it was measured

Every system was given the same material and scored by the same rules. The steps below condense the protocol and the scoring rules, and we share them before any pilot so your result is scored the same way.

  1. Source material

    Real multi-dialect Arabic business meetings, with human-verified reference transcripts prepared for every scored segment.

  2. Same audio for every system

    All seven systems received the same source audio, pre-segmented by the M3 capture workflow into participant-separated segments. External systems were scored on transcription only; their speaker diarization, the step that splits audio by who is speaking, was not used or evaluated. Every competitor figure on this page was produced by running the vendor's API on the evaluation audio; none is taken from a vendor publication.

  3. Configurations

    Google Cloud Speech-to-Text V2 as chirp_3 with ar-XA; Google Cloud Speech-to-Text V1 as latest_long with ar-SA; ElevenLabs as scribe_v1 with language_code=ara; Cohere, Deepgram Nova-3 and Sonix.ai as named, with no further model string given. M3 as the complete meeting system with participant-linked audio channels from its capture workflow.

  4. Metric

    CER = (substitutions + deletions + insertions) divided by the characters in the human-verified reference, after normalisation. Character accuracy is 100 minus CER.

  5. Equivalences

    A dialect spelling and its Modern Standard Arabic form count the same when the meaning is the same. An English word counts whether written in Latin letters, phonetically in Arabic script, or as an equivalent Arabic word. Numbers are compared by value (15, ١٥ and خمسة عشر are the same) and stay in the scored content, as do English business terms.

  6. Ignored

    Punctuation entirely; Latin case; Arabic letter variants such as the forms of alef are normalised; whitespace, except that a missing space that fuses two words counts as one error.

  7. Remaining differences

    Every remaining missing, added or substituted character is an error. Only segments with a verified reference enter the CER. An empty output against a speech-bearing reference counts as deletions; extra unsupported output counts as insertions.

  8. Qualitative review

    Reviewers recorded content omission, names and English terminology, filler-like insertions, unexpected language changes, repeated punctuation and repetition loops per system. No separate aggregate rates were calculated.

Evidence map

ClaimScopeStatus
M3 1.62% CER versus 2.26% for Google chirp_3, 4.93% for Cohere, 5.01% for ElevenLabs Scribe, 6.09% for Deepgram Nova-3, 9.01% for Google latest_long and 9.8% for Sonix; relative reductions of approximately 28%, 67%, 68%, 73%, 82% and 83%.The same human-verified evaluation set of real meetings, same reference, same scoring rules.Measured.
Vendor-specific behaviour as summarised in the matrix (name errors, segment-start omissions, filler-like vocalisations, a language or script switch, non-speech rejection, strong English-term handling, excessive sentence-final punctuation).Qualitative review of the same evaluation set.Observed qualitatively; separate aggregate rates not calculated.
Platform-provided speaker attribution remained linked throughout every evaluated meeting.M3 participant-channel workflow; external systems evaluated for transcription only.Measured for the M3 participant-channel workflow.
Approximately 80x batch throughput; about 45 seconds per hour of meeting audio.M3 meeting transcription workflow for a one-hour meeting on an L4 GPU.Team-reported; batch throughput, not live latency.
No external transcription call.Evaluated transcription path.Verified in the evaluated deployment.

Limitations

  • One evaluation set. Every CER figure comes from a single human-verified evaluation set of real meetings. It shows how the seven systems compared on that material, not a universal rate for any of them. Your meetings, dialect mix and microphones are the benchmark that matters.
  • A snapshot in time. The compared APIs were called in September 2026 with the configurations listed under "How it was measured". Vendors update their models, so the figures describe those versions on that date.
  • A small margin at the top. The gap to Google chirp_3 is 0.64 percentage points, and on several example segments chirp_3 is as complete as M3.
  • Throughput is team-reported. The 45 seconds per hour and 80x figures come from the team for one hardware configuration and are not part of the scored benchmark. Live latency was not measured.
  • Qualitative observations carry no aggregate rates. The matrix levels follow one rule applied to the wording of the reviewer notes, shown with their phrases so they can be challenged.
  • External speaker diarization was not evaluated. Every external system received participant-separated audio prepared by M3. The comparison isolates transcription and says nothing about how those systems attribute speakers on mixed audio. The M3 speaker-attribution result is a coverage statement; no speaker-attributed error rate is reported.
  • M3 is not perfect. The examples show hesitation sounds transcribed where the reference has none, and words added that the reference does not contain.
  • Next steps. The next evaluation phase expands testing across additional organisations, meeting platforms, dialects, acoustic conditions and customer-specific terminology. You can evaluate M3 on your own meetings now.
Full results, all seven systems
Full results, all seven systems. Measured on the same human-verified meetings, same reference, same scoring rules; see How it was measured. best in column
RankSystemProtocol detailCER (%), lower is betterCharacter accuracy (%)M3 relative reduction, approximate (%)Absolute reduction, points
1Complete M3 meeting systemComplete M3 meeting system; participant-linked audio channels supplied by the M3 capture workflow1.62 (best in column)98.38 (best in column)baselinebaseline
2Google Cloud Speech-to-Text V2 (chirp_3)chirp_3, ar-XA2.2697.74280.64
3CohereCohere (no model string given in the white paper)4.9395.07673.31
4ElevenLabs Scribe v1scribe_v1, language_code=ara5.0194.99683.39
5Deepgram Nova-3Deepgram Nova-3 (no further model string given)6.0993.91734.47
6Google Cloud Speech-to-Text V1 (latest_long)latest_long, ar-SA9.0190.99827.39
7Sonix.aiSonix.ai (no model string given)9.890.283 (best in column)8.18 (best in column)

The canonical report lives on lisan.com. Cite that page.

Cite this page

Lisan Research (2026). M3: Arabic meeting transcription, measured on real meetings. Evaluation Edition, September 2026. Lisan, 15 September 2026. https://lisan.com/company/research/m3/
BibTeX
@techreport{lisan2026m3,
  author = {Lisan Research},
  title = {M3: Arabic meeting transcription, measured on real meetings},
  institution = {Lisan},
  year = {2026},
  month = {September},
  type = {Evaluation report},
  note = {Evaluation Edition, September 2026. Published 15 September 2026},
  url = {https://lisan.com/company/research/m3/},
}

Run M3 on your own meetings

A pilot runs M3 on a sample of your meetings against a human-verified reference, with the scoring rules shared before the run, so your result is measured the same way this report was. Or start free and bring it to your next meeting.

Frequently asked questions

What does a 1.62% character error rate mean in practice?

Roughly 16 wrong characters in every 1,000, after the published equivalences (dialect and standard spellings, Latin and phonetic English, and numeric forms count the same). It was measured on real multi-dialect Arabic business meetings with human-verified references, on the same material as the six compared systems. It is a result on one evaluation set, not a guarantee for every room.

Is this the word error rate figure that used to be on this site?

No. That figure is superseded. This evaluation reports character error rate against six named systems with published scoring rules, and it replaces every earlier MeetriX accuracy figure on this site.

Was the evaluation independently verified?

No. Lisan Research ran it. The protocol, scoring rules, examples and evidence map are published so the scoring can be reproduced, and a pilot on your own meetings is the way to check it on your material.

Does MeetriX send our audio to a cloud speech API?

In the evaluated deployment the transcription path called no external transcription service (verified in the evaluated deployment). As a product capability, M3 can run within customer-controlled infrastructure so audio, transcripts and summaries stay inside your data-governance boundary; deployment options are on the security and on-prem page.

How do we run a pilot?

Book a pilot or write to support@lisan.com. We agree a sample of your meetings, prepare a human-verified reference, share the scoring rules before the run, and report the result the same way this page does.