Announcement
Introducing M3.
The Arabic meeting transcription system inside MeetriX now has a published evaluation: 1.62% character error rate on real multi-dialect Arabic meetings, the lowest of the seven systems compared, with every scoring rule and example on the page.
M3, the system that produces every MeetriX transcript of an Arabic meeting, recorded a 1.62% character error rate on real multi-dialect Arabic business meetings, the lowest of the seven systems compared. Today we are publishing the first evaluation report on it, and this post is the short version. M3 is the part of the product that listens to who said what, in whichever dialect they said it, with the English terms that arrive mid-sentence, and writes it down.
What we measured
We took real multi-dialect Arabic business meetings, prepared human-verified reference transcripts, and ran the same audio through seven systems: M3 and six commercial speech APIs, Google chirp_3, Cohere, ElevenLabs Scribe, Deepgram Nova-3, Google latest_long and Sonix. Every output was scored against the same reference with the same rules. The metric is character error rate: every character that is wrong, missing or extra, divided by the characters in the reference, after a published set of equivalences. A dialect spelling and its standard form count the same. An English term counts whether it is written in Latin letters or phonetically in Arabic script. Numbers are compared by value. Punctuation is ignored.
M3 recorded 1.62%, which means roughly 16 wrong characters in every 1,000. Google chirp_3 followed at 2.26%. Then Cohere at 4.93%, ElevenLabs Scribe at 5.01%, Deepgram Nova-3 at 6.09%, Google latest_long at 9.01% and Sonix at 9.8%. That is a measured result on one evaluation set, and the report says so.
M3 records the lowest character error rate of seven systems on the same human-verified meetings
Lower is better. Character error rate (CER) after defined lexical normalisation.
One example
Numbers hide what a failure looks like, so the report shows transcripts side by side. In a daily stand-up full of developer vocabulary, the human-verified reference keeps SDK, rendering, highlights, editor, configuration and Chrome extension in Latin script, and so does M3. Google chirp_3 transcribed the passage completely but phonetically, in Arabic letters, Cohere did the same and garbled the ending, and Sonix dropped the partner name twice. A phonetic spelling is not an error under the scoring rules; the difference is what a search across the archive will find later. The full page has four more, among them a hesitant status update in which the VAT term survives in M3 while several systems garble or drop it, a greeting shorter than a second that two systems returned nothing for, and a quality update with three spoken figures that only three systems kept intact.
What the other systems did well
Google chirp_3 handled English technical terminology strongly within Arabic speech and showed low content loss, and on several example segments it is as complete as M3, with the difference limited to phonetic rather than Latin spellings. ElevenLabs Scribe preserved English terminology particularly well. Cohere was highly accurate on Classical Arabic. All seven systems supported the Arabic dialects present in the evaluated meetings. The report also lists M3's own mistakes, including hesitation sounds transcribed where the reference has none and words added that the reference does not contain.
What it means for your meetings
The transcription path in the evaluated deployment called no external transcription service. MeetriX can also run on your own servers with the speech engine inside your network. The M3 workflow processes a one-hour meeting in approximately 45 seconds on an L4 GPU, about 80x real time; that figure is team-reported and describes batch processing of recorded audio, not live latency. The transcript stays linked to its speakers and timestamps, and the summary, minutes and action items are built from it.
Older MeetriX pages quoted a word error rate against an unnamed cloud provider. That figure is superseded: from today the site cites this evaluation, with the systems named and the rules published. Read the full results, see how the engine handles Arabic and English in one meeting, or book a pilot and we will score M3 on your own meetings the same way.
Stay in the conversation. MeetriX takes the notes.
It joins the call, transcribes who said what in Arabic or English, and sends the summary and action items before you are back at your desk. 600 free minutes when you connect your calendar.
Arabic & English · 32 Arabic dialects · No credit card required