Skip to main content
Arabic + English · on-premise or private cloud

Every meeting, distilled to what matters

Khulasa sends a bot to your Google Meet, Zoom, or Teams call, records it, and returns a clean summary, decisions, and action items — in Arabic and English. Available as a private cloud or on-premise deployment, for teams that need it inside their own infrastructure.

No credit card required

Encrypted in transit and at rest, tenant-isolated, and never used to train models unless you opt in — with on-premise and private-cloud options for regulated teams.

meet.google.com/abc-defg-hij
Recording

Summary

Action items

Built for bilingual teams

From join to summary, every step is automated and secure.

Joins & records automatically

Schedule a meeting and the bot joins on time across Google Meet, Zoom, and Teams.

Bilingual transcription

Transcribes Arabic and English in one sentence, keeping English terms in Latin script instead of transliterating them.

Summaries that decide

Get the summary, decisions, risks, and owner-assigned action items — not just a wall of text.

Roles & sharing

Organization roles, workspace groups, and per-recording sharing keep access exactly where it should be.

Deploy on your own infrastructure

Deploy on-premise or as a private cloud, for regulated teams that need meeting data to stay inside their own infrastructure. The speech and language models can run on your own hardware.

How it works

01

Paste a meeting link

Or schedule it — one-off or recurring.

02

The bot records

It joins, captures audio or video, and uploads securely.

03

Get your analysis

Summary, decisions, and action items, ready to share.

Benchmarked · August 2026

Measured against other systems

We ran our self-hosted engine and every Arabic-focused competitor we could reach through a public API over the same real, bilingual meeting audio, and scored them all the same way. Here is the whole result — including the Arabic column, where three of the four competing configurations beat us.

35.14

Overall word error rate

The lowest overall error rate in this comparison — but we don't claim a win. Two of the four paired intervals span zero, so against the leading vendors we report parity rather than a ranking.

+16.3+34.7points

English words inside Arabic speech

More accurate on the English words spoken mid-Arabic-sentence than each of the four competing configurations, one by one — paired intervals at 95% confidence, all excluding zero. This is the one axis where we claim the lead.

>3×

Arabic–English switch points kept

When someone switches between Arabic and English mid-sentence, more of those moments survive into the transcript — over three times as many as any of the three Speechmatics configurations kept. The error counted is a script collapse: an English word written in Arabic letters.

Every figure here describes the self-hosted configuration, running on your own hardware. We ran this comparison ourselves on 80 code-switched utterances of real UN ESCWA meeting audio. It is not a measurement of our hosted tiers.

One dataset: real UN ESCWA meeting audio — 80 utterances, every one code-switched.

Swipe the table to see every column.

Benchmark comparison of speech recognition systems on identical audio
SystemOverall WERLower is betterArabic wordsLower is betterEnglish wordsLower is betterSwitch points keptHigher is better
Khulasa — self-hosted35.1437.1629.0839.0%
Speechmaticsstandard Arabic/English37.2934.6945.4011.0%
Speechmatics melia-1code-switch model38.6030.2663.8112.1%
Speechmaticsenhanced Arabic/English40.1832.3763.187.4%
A leading Arabic cloud vendorcode-switch model43.8540.3954.6025.0%

We do not claim to win this column. Two of the four paired intervals span zero, so at this sample size we report parity with the leading vendors rather than a ranking. We do not win the Arabic column either — all three Speechmatics configurations are ahead of us there. The English column is the one where every interval excludes zero.

Read the full methodology

Why the English column decides an Arabic meeting

An Arabic meeting is not an Arabic-only meeting. The product names, the tools, the metrics, the deadlines — the words an action item is actually about — arrive in English, mid-sentence, in Latin script.

Overall word error rate is dominated by the Arabic tokens, because in this audio there are more of them. A system can therefore score well overall while writing the English words in Arabic letters: the aggregate barely moves, and the English word is gone from the transcript. Nothing downstream can recover a word that was never written down.

So read the table as a trade, not a sweep. All three Speechmatics configurations transcribe the Arabic words more accurately than this one does; on the English words this one is far ahead of every system here; the overall rate comes out level. In a bilingual meeting that is the trade we would take — and it is the axis we measured and published.

Across the four competing configurations the collapse count runs from 38 to 164, the highest being Speechmatics melia-1; this configuration collapsed 53, which is not the lowest figure in the comparison.

The model is open — how we run it is ours

The speech model is Whisper large-v3 — from OpenAI, MIT licence, open weights. We did not train it, and we lead with that. What is ours is everything around it: the model runs at the published reference configuration its authors describe rather than at the library default, language is decided once per recording instead of once per turn, and the deployment is on-premise, so the audio stays on your own hardware. Then we measured it on bilingual meetings and published what came back.

Whisper large-v3OpenAIMIT

How this was measured

We ran this comparison ourselves. Every system transcribed identical audio and was scored by the same normalizer. We report corpus-level word error rate with a paired per-utterance bootstrap — 2,000 draws at 95% confidence — and where an interval spans zero we call the systems not distinguishable rather than naming a winner. The audio is 80 utterances of real UN ESCWA meeting audio, every one code-switched, selected by a hash of the utterance identifier rather than chosen by us. Every figure describes the self-hosted configuration of Khulasa, the one that runs on your own hardware, on-premise. Measured August 2026.

English words inside Arabic speechPaired intervals at 95% confidence, all excluding zero — most conservative lower bound+8.8

Arabic–English switch points kept39.0%vs7.4%12.1%(Speechmatics)

Corpora

Used under their published licences. Only the ESCWA audio carries the figures in the table; the others are named because the limits below draw on them.

What we do not claim

  • These are public research corpora, not customer meetings. A real meeting is harder than all of them, and no customer audio has ever been measured.
  • One corpus of 80 utterances is a small sample, which is why we publish intervals rather than rankings.
  • We make no per-dialect claim — there are too few clips per dialect for this data to resolve one, and Egyptian Arabic is not covered at all.
  • On dialectal Arabic that is not meeting audio, the leading vendors are ahead of this configuration — by around 14 points on Saudi broadcast. We chose this configuration for bilingual meetings, and that is the trade we made.
  • Word error rate is not summary quality. A system can score well here and still produce a fluent, confident, wrong sentence, and nothing on this page measures the summary.
  • Every figure here describes the self-hosted, on-premise configuration — the one that runs on your own hardware. This is not a measurement of our hosted tiers.
  • On English-only meeting audio the leading vendor is ahead of this configuration — by around 4 to 6 points. The advantage we report is specific to English words spoken inside Arabic sentences.
  • Intella and Notah are not in this table. Neither exposes a public API we could run this audio through, so neither could be measured here.

Scope and trademarks

  • Not every system here is named. Where a name is not needed to identify the configuration measured, we describe it instead; the rules on comparative naming differ across the markets this page is read in. Every system was run, scored and reported the same way, and the label changes no figure.
  • Each competing system was run through its vendor's public API, in the configuration shown in the table, on the audio described above, in August 2026. Our own row is the self-hosted engine at its shipped configuration.
  • Every system in this table is a bilingual Arabic-and-English configuration; no Arabic-only result is reported here. Where a vendor offers more than one such configuration we ran more than one: three Speechmatics configurations were run on this audio and all three are shown, and for the vendor described rather than named we ran every model its API exposes and report its best result.
  • These figures describe those configurations, on that audio, on that date. Vendors update their models, and the same run repeated later may not reproduce them.
  • Each corpus above is used under the licence printed beside it, and several are restricted to non-commercial research. What we publish are measurements computed from that audio; we redistribute no audio and no reference transcript, in whole or in part.
  • Product and company names are the trademarks of their respective owners. Khulasa is not affiliated with, sponsored by or endorsed by any of them, and no vendor reviewed or approved these results.
  • If a vendor believes a configuration shown here is not their best available, tell us and we will re-run the comparison.
Security & privacy

Your meetings stay yours

Protection built in from the moment the bot joins to long-term storage.

Encrypted end to end

Recordings and transcripts are encrypted in transit (TLS) and at rest.

Granular access control

Organization roles, workspace groups, and per-recording sharing decide exactly who can see what.

Tenant isolation

Every organization's data is strictly separated — your recordings never mix with anyone else's.

Training is off by default

Your meeting content is processed to produce your analysis and to run and secure the service. We do not use it to train AI models — training is opt-in, and it is off by default.

Simple, seat-based pricing

Start free. Add seats and power as your team grows.

Free

Free

  • Up to 3 seats
  • 3 hours of recording / user / month
See what's included

Team

$20 / seat / mo

  • Add seats as you grow
  • 30 hours of recording / user / month
See what's included
Most popular

Business

$34 / seat / mo

  • Add seats as you grow
  • 120 hours of recording / user / month
See what's included
For organizations

Private Cloud

Contact us

A dedicated, isolated cloud deployment run for your organization.

For organizations

On-Premise

Contact us

For regulated, government, and data-sovereignty-bound teams. Deployed and run on your own servers.

Turn your next meeting into decisions

Create an organization in under a minute.

Create your account