Every meeting, distilled to what matters
Built for bilingual teams
From join to summary, every step is automated and secure.
Joins & records automatically
Schedule a meeting and the bot joins on time across Google Meet, Zoom, and Teams.
Bilingual transcription
Transcribes Arabic and English in one sentence, keeping English terms in Latin script instead of transliterating them.
Summaries that decide
Get the summary, decisions, risks, and owner-assigned action items — not just a wall of text.
Roles & sharing
Organization roles, workspace groups, and per-recording sharing keep access exactly where it should be.
Deploy on your own infrastructure
Deploy on-premise or as a private cloud, for regulated teams that need meeting data to stay inside their own infrastructure. The speech and language models can run on your own hardware.
How it works
Paste a meeting link
Or schedule it — one-off or recurring.
The bot records
It joins, captures audio or video, and uploads securely.
Get your analysis
Summary, decisions, and action items, ready to share.
Measured against other systems
We ran our self-hosted engine and every Arabic-focused competitor we could reach through a public API over the same real, bilingual meeting audio, and scored them all the same way. Here is the whole result — including the Arabic column, where three of the four competing configurations beat us.
35.14
Overall word error rate
The lowest overall error rate in this comparison — but we don't claim a win. Two of the four paired intervals span zero, so against the leading vendors we report parity rather than a ranking.
+16.3 … +34.7points
English words inside Arabic speech
More accurate on the English words spoken mid-Arabic-sentence than each of the four competing configurations, one by one — paired intervals at 95% confidence, all excluding zero. This is the one axis where we claim the lead.
>3×
Arabic–English switch points kept
When someone switches between Arabic and English mid-sentence, more of those moments survive into the transcript — over three times as many as any of the three Speechmatics configurations kept. The error counted is a script collapse: an English word written in Arabic letters.
Every figure here describes the self-hosted configuration, running on your own hardware. We ran this comparison ourselves on 80 code-switched utterances of real UN ESCWA meeting audio. It is not a measurement of our hosted tiers.
One dataset: real UN ESCWA meeting audio — 80 utterances, every one code-switched.
Swipe the table to see every column.
| System | Overall WER†Lower is better | Arabic wordsLower is better | English wordsLower is better | Switch points keptHigher is better |
|---|---|---|---|---|
| Khulasa — self-hosted | 35.14 | 37.16 | 29.08 | 39.0% |
| Speechmaticsstandard Arabic/English | 37.29 | 34.69 | 45.40 | 11.0% |
| Speechmatics melia-1code-switch model | 38.60 | 30.26 | 63.81 | 12.1% |
| Speechmaticsenhanced Arabic/English | 40.18 | 32.37 | 63.18 | 7.4% |
| A leading Arabic cloud vendorcode-switch model | 43.85 | 40.39 | 54.60 | 25.0% |
We do not claim to win this column. Two of the four paired intervals span zero, so at this sample size we report parity with the leading vendors rather than a ranking. We do not win the Arabic column either — all three Speechmatics configurations are ahead of us there. The English column is the one where every interval excludes zero.
Read the full methodologyHide the method
Why the English column decides an Arabic meeting
An Arabic meeting is not an Arabic-only meeting. The product names, the tools, the metrics, the deadlines — the words an action item is actually about — arrive in English, mid-sentence, in Latin script.
Overall word error rate is dominated by the Arabic tokens, because in this audio there are more of them. A system can therefore score well overall while writing the English words in Arabic letters: the aggregate barely moves, and the English word is gone from the transcript. Nothing downstream can recover a word that was never written down.
So read the table as a trade, not a sweep. All three Speechmatics configurations transcribe the Arabic words more accurately than this one does; on the English words this one is far ahead of every system here; the overall rate comes out level. In a bilingual meeting that is the trade we would take — and it is the axis we measured and published.
Across the four competing configurations the collapse count runs from 38 to 164, the highest being Speechmatics melia-1; this configuration collapsed 53, which is not the lowest figure in the comparison.
The model is open — how we run it is ours
The speech model is Whisper large-v3 — from OpenAI, MIT licence, open weights. We did not train it, and we lead with that. What is ours is everything around it: the model runs at the published reference configuration its authors describe rather than at the library default, language is decided once per recording instead of once per turn, and the deployment is on-premise, so the audio stays on your own hardware. Then we measured it on bilingual meetings and published what came back.
Whisper large-v3OpenAIMIT
How this was measured
We ran this comparison ourselves. Every system transcribed identical audio and was scored by the same normalizer. We report corpus-level word error rate with a paired per-utterance bootstrap — 2,000 draws at 95% confidence — and where an interval spans zero we call the systems not distinguishable rather than naming a winner. The audio is 80 utterances of real UN ESCWA meeting audio, every one code-switched, selected by a hash of the utterance identifier rather than chosen by us. Every figure describes the self-hosted configuration of Khulasa, the one that runs on your own hardware, on-premise. Measured August 2026.
English words inside Arabic speechPaired intervals at 95% confidence, all excluding zero — most conservative lower bound≥ +8.8
Arabic–English switch points kept39.0%vs7.4%–12.1%(Speechmatics)
Corpora
Used under their published licences. Only the ESCWA audio carries the figures in the table; the others are named because the limits below draw on them.
- UN ESCWA meeting audio — the corpus scored aboveCC BY-NC 4.0
- Mixat — Emirati Arabic/English podcastCC BY-NC-SA 4.0
- SADA — Saudi broadcast audioCC BY-NC-SA 4.0
- AMI — English meeting audio, far-field and close-talkCC BY 4.0
What we do not claim
- These are public research corpora, not customer meetings. A real meeting is harder than all of them, and no customer audio has ever been measured.
- One corpus of 80 utterances is a small sample, which is why we publish intervals rather than rankings.
- We make no per-dialect claim — there are too few clips per dialect for this data to resolve one, and Egyptian Arabic is not covered at all.
- On dialectal Arabic that is not meeting audio, the leading vendors are ahead of this configuration — by around 14 points on Saudi broadcast. We chose this configuration for bilingual meetings, and that is the trade we made.
- Word error rate is not summary quality. A system can score well here and still produce a fluent, confident, wrong sentence, and nothing on this page measures the summary.
- Every figure here describes the self-hosted, on-premise configuration — the one that runs on your own hardware. This is not a measurement of our hosted tiers.
- On English-only meeting audio the leading vendor is ahead of this configuration — by around 4 to 6 points. The advantage we report is specific to English words spoken inside Arabic sentences.
- Intella and Notah are not in this table. Neither exposes a public API we could run this audio through, so neither could be measured here.
Scope and trademarks
- Not every system here is named. Where a name is not needed to identify the configuration measured, we describe it instead; the rules on comparative naming differ across the markets this page is read in. Every system was run, scored and reported the same way, and the label changes no figure.
- Each competing system was run through its vendor's public API, in the configuration shown in the table, on the audio described above, in August 2026. Our own row is the self-hosted engine at its shipped configuration.
- Every system in this table is a bilingual Arabic-and-English configuration; no Arabic-only result is reported here. Where a vendor offers more than one such configuration we ran more than one: three Speechmatics configurations were run on this audio and all three are shown, and for the vendor described rather than named we ran every model its API exposes and report its best result.
- These figures describe those configurations, on that audio, on that date. Vendors update their models, and the same run repeated later may not reproduce them.
- Each corpus above is used under the licence printed beside it, and several are restricted to non-commercial research. What we publish are measurements computed from that audio; we redistribute no audio and no reference transcript, in whole or in part.
- Product and company names are the trademarks of their respective owners. Khulasa is not affiliated with, sponsored by or endorsed by any of them, and no vendor reviewed or approved these results.
- If a vendor believes a configuration shown here is not their best available, tell us and we will re-run the comparison.
Your meetings stay yours
Protection built in from the moment the bot joins to long-term storage.
Encrypted end to end
Recordings and transcripts are encrypted in transit (TLS) and at rest.
Granular access control
Organization roles, workspace groups, and per-recording sharing decide exactly who can see what.
Tenant isolation
Every organization's data is strictly separated — your recordings never mix with anyone else's.
Training is off by default
Your meeting content is processed to produce your analysis and to run and secure the service. We do not use it to train AI models — training is opt-in, and it is off by default.
Simple, seat-based pricing
Start free. Add seats and power as your team grows.
Business
$34 / seat / mo
- Add seats as you grow
- 120 hours of recording / user / month
Private Cloud
Contact us
A dedicated, isolated cloud deployment run for your organization.
On-Premise
Contact us
For regulated, government, and data-sovereignty-bound teams. Deployed and run on your own servers.