Research
Same meetings, same references,
same rules for every system.
What we measured, how we measured it, and what the results do and do not support. Every report publishes its scoring rules, its examples and its limitations.
How MeetriX evaluates
Every system in a MeetriX evaluation gets the same meetings, the same human-verified references and the same scoring rules. The rules are published before the numbers: which spellings count as equivalent, how numbers are compared, what an empty output costs. Competitor systems are named, with the model version wherever the vendor exposes one, and the configuration used is listed under the methodology; the report says plainly where they did well. Figures that were reported by the team rather than measured, such as throughput, are labelled that way, and qualitative observations are never turned into rates that were not calculated. Test-set sizes are not used as headlines; the material is described as what it is, real multi-dialect Arabic business meetings with human-verified reference transcripts.
The canonical version of every report lives on the Lisan research hub, with the citation record. The pages here add the product view: what a result means for a meeting record, and how to run the same evaluation on your own meetings.
Run the same evaluation on your meetings
A pilot scores M3 on a sample of your meetings against a human-verified reference, with the rules shared before the run.