Falcon ASR: Arabic and English speech recognition, with a headset falcon and the greetings هلا والله and Hello world.

TECHNOLOGY INNOVATION INSTITUTE

Introducing Falcon ASR

20.92%Arabic WER
1.6BParameters
22.73%Emirati WER · TII evaluation

We’re introducing Falcon-ASR, our 1.6 billion parameter speech recognition model for Arabic, with a particular focus on the Emirati dialect. Developed at the Technology Innovation Institute (TII) in Abu Dhabi, it also supports English, French, Spanish and Portuguese.

In our evaluation, Falcon-ASR achieved an average word error rate of 20.92% across six Arabic test sets, compared with the best published result of 23.17% in the leaderboard snapshot we used. On our internal Emirati evaluation, it recorded the lowest word and character error rates among the systems we compared.

We also support word-level timestamps for transcriptions, linking each transcribed word to its position in the audio.

You can try Falcon-ASR in our Hugging Face Demo.

Recognising spoken Arabic

Arabic speech varies by region, speaker and setting. A model that handles a formal news broadcast may still struggle with a conversation in Emirati or with speech recorded over a phone line. Dialectal Arabic also has fewer transcribed resources than Modern Standard Arabic (MSA), which makes training and evaluation harder.

We trained Falcon-ASR on Emirati, MSA, other Gulf and Arabic dialects, and English. Our aim is to transcribe the words people use in everyday speech, including dialectal forms and changes between languages.

Arabic benchmark results

The Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, ranks systems by the equal-weight average WER across six test sets. It also reports character error rate (CER). Lower values are better for both metrics. Our Falcon-ASR evaluation follows this protocol.

ModelParametersAvg WER (%)Avg CER (%)
Falcon-ASR1.6B20.928.79
Audar-ASR-V1-Turbo2.35B23.179.23
Cohere Transcribe Arabic (07-2026)2.0B25.8711.80
omniASR LLM 7B7.0B28.3212.52

WER = Word Error Rate; CER = Character Error Rate. A lower value indicates better performance.

We evaluated Falcon-ASR on the same six benchmarks using the leaderboard’s pinned manifests. Competitor figures are the published leaderboard averages checked on 30 September 2026. Falcon-ASR’s average WER is 2.25 percentage points better than the best published result in that snapshot.

Evaluating Emirati speech

Public evaluation data already includes Emirati: Casablanca has a UAE subset. We complement that coverage with an internal evaluation of additional Emirati and Gulf speech, using held-out recordings and human-validated transcripts to assess transcription accuracy beyond the public UAE subset.

In our internal Emirati evaluation, Falcon-ASR achieved 22.73% WER and 10.19% CER:

ModelParametersWER (%)CER (%)
Falcon-ASR1.6B22.7310.19
Qwen3-Omni-30B-A3B-Instruct30.0B (3.0B active)26.8012.72
Audar-ASR-V1-Turbo2.35B27.8913.75
Cohere Transcribe Arabic (07-2026)2.0B31.0518.07
Qwen3-ASR-1.7B-hf2.0B31.5213.35
Audar-ASR-V1-Flash0.78B32.8715.36

Falcon-ASR has the lowest WER and CER among the systems compared here. Its WER is 4.07 percentage points below Qwen3-Omni, the next best result. The results show improved transcription accuracy at both the word and character level on this evaluation.

Training for different recording conditions

We included background noise, overlapping speech, music, room reverberation and telephony effects, as well as variations in speed and pitch. We applied the same treatment to Emirati recordings, exposing the model to a range of conditions it may encounter in meetings, calls and other everyday recordings.

English and other languages

Falcon-ASR also transcribes English with the same model weights. In our evaluation on the seven public English test sets used by the Hugging Face Open ASR Leaderboard, it achieved a mean WER of 5.74%.

Test setWER (%)
LibriSpeech clean1.75
LibriSpeech other4.21
SPGISpeech2.02
VoxPopuli3.87
GigaSpeech8.15
AMI8.33
Earnings-2211.86

The model also supports French, Spanish and Portuguese. All five languages use the same weights, without requiring a language flag. The output is a transcript in the language spoken.

Model foundation

Falcon-ASR builds on our Falcon3-Audio work. The architecture and training approach for Falcon3-Audio are described in Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data.

Try Falcon ASR

Our Hugging Face Demo Space lets you try Falcon-ASR and explore its transcription capabilities. API access and native applications are planned. We invite you to try the Demo with your own recordings.