--- license: mit language: - ar - en tags: - speech - asr - automatic-speech-recognition - ctc - conformer - egyptian-arabic - code-switching - arabic - audio - pytorch - kenlm - language-model - streaming library_name: metro-asr pipeline_tag: automatic-speech-recognition datasets: - AlaaSamir/custom-egy-tts - OmarAhmedSobhy/egyption-with-emotion-dataset - MightyStudent/Egyptian-ASR-MGB-3 - MAdel121/arabic-egy-cleaned - MAdel121/Continuation-egy-for-ultravox-v1 - Raniahossam33/Egyptian_TTS3RS - ahmedbasemdev/egyptain-tts-dataset - MohamedRashad/arabic-english-code-switching - librispeech_asr metrics: - wer - cer model-index: - name: Metro-ASR Small results: - task: type: automatic-speech-recognition name: Speech Recognition dataset: type: custom name: Egyptian Arabic + Code-Switching Test Set config: all split: test metrics: - type: wer value: 46.85 name: WER (All) - type: cer value: 28.41 name: CER (All) - type: wer value: 37.24 name: WER (Arabic) - type: cer value: 17.45 name: CER (Arabic) - type: wer value: 36.32 name: WER (Code-Switching) - type: cer value: 17.44 name: CER (Code-Switching) ---

Metro-ASR Small

Non-autoregressive CTC speech recognition for Egyptian Arabic and Arabic–English
code-switching, with a detachable n-gram language head you can retrain on text alone.

GitHub PyPI Space Evaluation report

## What this is Metro-ASR separates the two things a speech recogniser has to know — **what the audio sounds like** and **what the words are likely to be** — into two artefacts trained and shipped separately. This repository holds the first one: **Metro-Small**, a 61.6M-parameter Conformer acoustic model trained with CTC. It is non-autoregressive — one forward pass turns an utterance into a matrix of per-frame log-probabilities, with no decoder loop and no dependence on previously emitted tokens — which is why it runs at 40–55× real time on a laptop CPU with no GPU involved anywhere in this card's examples. The second artefact, the language model, is an n-gram over text. It never sees audio, trains in minutes on a laptop, and plugs into the decoder at run time — including domain-specific variants you build yourself from nothing but text (see [The language head](#the-language-head) below). The general-purpose one is included in this repository as `lm_5gram.bin`. For the full architecture writeup, training-from-scratch instructions, and how to build your own domain language head, see the **[GitHub repository](https://github.com/MohammedAly22/metro-asr)**. For real transcripts, audio players, and a measured comparison of greedy vs. beam+LM decoding across three interchangeable language heads, see the **[interactive evaluation report](https://mohammedaly22.github.io/metro-asr/)**.

Metro-ASR architecture

--- ## Architecture Twelve identical Conformer-style blocks, each a Macaron sandwich of two half-weighted feed-forward networks around an attention module and a convolution module:

Inside one Metro block

| Component | Details | Why | |---|---|---| | Encoder | Conformer, 12 layers, d_model=384, 6 heads | — | | Position encoding | RoPE (rotary), bias-free Q/K projections | attention depends on relative offset, so it doesn't break on utterances longer than any seen in training | | Feed-forward | SwiGLU, expansion 3×, applied twice at half weight | a learned gate beats a plain ReLU/GELU FFN at equal parameter count | | Normalization | RMSNorm, pre-norm | cheaper than LayerNorm, keeps gradients well-behaved as depth grows | | Convolution | SE-gated depthwise separable, kernel 31 | local context (~1.24 s) to complement attention's global view | | Regularization | Stochastic depth, rate 0.05 | deeper layers dropped more often during training | | Auxiliary loss | Intermediate CTC at layer 6, weight 0.3 | mid-stack layers get gradient directly; discarded at inference | | Tokenizer | BPE (SentencePiece), vocab 5,000 | trained on a deliberately balanced Arabic/English corpus so English words survive as whole tokens | | Decoding | CTC greedy, or beam search + KenLM | see below | | Parameters | 61,586,320 (61.6M) | — | **Frame rate.** 16 kHz audio → 80-bin log-Mel (100 fps) → Conv2D ×4 subsampling → 25 fps through the encoder and the CTC head. Subsampling by 4 before the first block cuts attention's quadratic cost 16× before a single block runs; one output token covers 40 ms of audio. --- ## The language head

The detachable language head

CTC's per-frame independence produces a specific, recognisable error pattern: doubled syllables, dropped affixes, malformed English fragments — the model heard correctly and wrote something that isn't a word. A word-level n-gram model fixes this during beam search, because it knows which *sequences* are plausible, without ever having heard a single second of audio:

Greedy vs. beam search with the language head

The language head is a separate file with no learned interaction with the acoustic weights. Swap it, and nothing about `model.pt` changes: ```python engine = MetroASREngine.from_pretrained("small") engine.load_lm("lm/medical_head.bin") # swap language heads at run time, print(engine.transcribe("call.wav", beam_search=True).text) # same acoustic weights throughout ``` **This is measured, not asserted.** The GitHub repo ships two extra language heads — technical and medical — built from real text (Egyptian medical chat/QA, Egyptian Arabic Wikipedia's technical articles, real Arabic-English code-switching text) plus synthesised domain-term carrier phrases, and decodes 11 real clips with all three heads against human references. Headline results from the **[full interactive report](https://mohammedaly22.github.io/metro-asr/)**: | Test set | Greedy | General head | Technical head | Medical head | |---|---:|---:|---:|---:| | Technical clips (WER) | 34.8% | 26.2% | **24.1%** | 26.2% | | Medical clips (WER) | 44.9% | **34.7%** | 38.8% | **34.7%** | | General speech (WER) | 25.6% | **24.7%** | 36.9% | 32.5% | Two things worth reading out of that table. The domain heads win in their own domain despite being 4-grams built from far less text than the general 5-gram — domain fit beats scale for this component. And the technical head actively *hurts* on general speech (36.9% vs. greedy's 25.6%) — a language head is a strong prior, and matching it to your traffic matters. (These numbers are from a small, hard 11-clip demo set — unscripted, overlapping speech, dense code-switching — and are not the same evaluation as the Performance section below; see the report for methodology.) Building your own head takes text and minutes, no GPU: ```bash python scripts/train_lm.py --corpus my_domain.txt --out lm/my_domain_4gram.arpa --order 4 ``` Full walkthrough — including a from-scratch corpus-building example for a technical and a medical domain — in the GitHub README's [Domain-specialised heads](https://github.com/MohammedAly22/metro-asr#domain-specialised-heads) section. --- ## Model variants

Metro-ASR family scaling

Three sizes share one block definition and one training recipe — only width, depth and vocabulary change. **Small (this repository) is the only one currently trained.** Medium and Large exist as configs in the GitHub repo with exact parameter counts, but no weights — see [Scaling to Medium and Large](https://github.com/MohammedAly22/metro-asr#scaling-to-medium-and-large) for what training them actually requires (data volume most of all). | | Params | d_model | Layers | BPE vocab | Status | |---|---:|---:|---:|---:|---| | **Small** | 61.6M | 384 | 12 | 5,000 | **Released — this repo** | | Medium | 247.4M | 512 | 24 | 8,000 | Config only | | Large | 747.8M | 768 | 32 | 16,000 | Config only | --- ## Performance Held-out test-set WER/CER (the numbers in this card's metadata): | Split | WER (%) | CER (%) | |---|---:|---:| | All | 46.85 | 28.41 | | Arabic only | 37.24 | 17.45 | | Code-switching | 36.32 | 17.44 | **Speed** — measured, Intel Core Ultra 7 155H, 4 CPU threads, PyTorch 2.13 CPU build, fp32, minimum of 15 runs after warm-up:

Measured CPU latency

| Audio length | Latency | RTF | Faster than real time | |---|---:|---:|---:| | 1 s | 37 ms | 0.037 | 27× | | 5 s | 99 ms | 0.020 | 51× | | **10 s** | **181 ms** | **0.018** | **55×** | | 30 s | 686 ms | 0.023 | 44× | Beam search with the 5-gram head adds roughly 10–250 ms per utterance depending on length (RTF ≈ 0.024 overall). Loading the 5.9 GB binary itself takes about 3.4 s, once, at startup. --- ## Usage ### Install ```bash pip install metro-asr # greedy decoding only pip install "metro-asr[lm]" # + KenLM beam search pip install -U "numpy>=2.0" # see note below ``` > [!WARNING] > `pyctcdecode`'s only PyPI release pins `numpy<2.0.0` in its own metadata, even though it runs > fine under numpy 2.x. Installing the `[lm]` extra will downgrade numpy to satisfy that — on an > environment that already had numpy 2.x with other packages built against it (Colab, most fresh > installs today), that breaks those packages with `numpy.dtype size changed`. The third line > above fixes it. If `[lm]` isn't installed at all, `lm_path="auto"` now degrades to greedy with > a warning rather than crashing engine construction. ### Quick start ```python from metro_asr import MetroASREngine engine = MetroASREngine.from_pretrained("small") # auto-downloads weights + tokenizer, caches locally result = engine.transcribe("audio.wav") print(result.text) ``` The 5.9 GB language model is **not** downloaded by this call. ### With the language head (beam search) ```python engine = MetroASREngine.from_pretrained("small", lm_path="auto") # also fetches lm_5gram.bin result = engine.transcribe("audio.wav", beam_search=True) print(result.text) ``` ### From a manual download If you've already run `snapshot_download` (or `git clone`d this repo) into a local directory, point `from_pretrained` at that directory instead — nothing gets re-downloaded, and it works fully offline: ```python from huggingface_hub import snapshot_download from metro_asr import MetroASREngine snapshot_download(repo_id="MohammedAly22/metro-asr-small", local_dir="checkpoints") engine = MetroASREngine.from_pretrained("checkpoints", lm_path="auto") ``` > [!TIP] > This is the fix for a common mistake: calling `from_pretrained("checkpoints")` used to be > interpreted as a HuggingFace repo id named literally "checkpoints" and fail with > *Repository Not Found*. Current versions check for an existing local directory first — update > if you hit that error. ### Without the package — loading the raw PyTorch model ```python import torch from metro_asr.utils.config import load_config from metro_asr.model.metro import MetroASR from metro_asr.model.tokenizer import build_tokenizer config = load_config("checkpoints/config.yaml") tokenizer = build_tokenizer(config, "checkpoints") # must run before MetroASR.from_config — model = MetroASR.from_config(config) # it fixes the CTC head's vocab size ckpt = torch.load("checkpoints/model.pt", map_location="cpu", weights_only=False) model.load_state_dict(ckpt["model_state_dict"]) model.eval() ``` ### Batch transcription ```python results = engine.transcribe_batch(["audio1.wav", "audio2.wav", "audio3.wav"]) for r in results: print(f"{r.text} (RTF={r.rtf:.4f})") ``` ### Streaming and serving `engine.transcribe_stream(chunk_generator)` yields incremental transcriptions from any audio generator, and `scripts/serve.py` in the GitHub repo wraps the same engine in a Flask REST API (`/transcribe`, `/transcribe/batch`, `/health`, `/info`). See the README's [Streaming](https://github.com/MohammedAly22/metro-asr#streaming) and [Serving](https://github.com/MohammedAly22/metro-asr#serving) sections, or the runnable [streaming_server.ipynb](https://github.com/MohammedAly22/metro-asr/blob/main/examples/streaming_server.ipynb) notebook. --- ## Files in this repository | File | Description | Size | |---|---|---:| | `model.pt` | Checkpoint — weights + AdamW optimizer state | 705 MB | | `config.yaml` | Model architecture configuration | <1 KB | | `bpe.model` | SentencePiece BPE tokenizer | 316 KB | | `bpe.vocab` | Human-readable vocabulary listing | 70 KB | | `lm_5gram.bin` | KenLM 5-gram general-purpose language head (optional) | 5.9 GB | Weights alone are 235 MB; `model.pt` is larger because it also carries optimizer state so training can be resumed from it. Strip that for deployment: ```python import torch ckpt = torch.load("model.pt", map_location="cpu", weights_only=False) torch.save({"model_state_dict": ckpt["model_state_dict"], "config": ckpt["config"]}, "model_inference.pt") ``` --- ## Training - **Audio data:** 130K+ clips from the 8 audio datasets listed in this card's metadata, plus additional Egyptian Arabic content, covering Arabic-only and Arabic-English code-switching speech. - **Text data (for the shipped language head):** ~1.9M Egyptian Arabic sentences plus Arabic-English code-switching text, upsampled to balance against the larger Arabic-only portion. - **Recipe:** CTC loss + 0.3-weighted auxiliary CTC at layer 6, AdamW (β = 0.9, 0.98), linear warmup into cosine decay, SpecAugment, speed perturbation (0.9×/1.0×/1.1×), 443K steps, batch size 32 with 4× gradient accumulation, bf16. - **Hardware:** single GPU. Full step-by-step instructions to reproduce this from scratch — tokenizer training, language model training, data preparation, the acoustic training loop, and how to scale the recipe to Medium/Large — are in the GitHub README's [Training from scratch](https://github.com/MohammedAly22/metro-asr#training-from-scratch) section. ### Fine-tuning this checkpoint Adapting to a new domain or accent starts from these weights, not from scratch — 10-20× lower learning rate, encoder frozen for the first few thousand steps so the CTC head adapts first: ```bash python scripts/finetune.py \ --checkpoint checkpoints/model.pt \ --tokenizer-dir checkpoints \ --dataset your/dataset-id \ --lr 5e-5 --max-steps 30000 --freeze-steps 3000 ``` Runnable end-to-end in [fine_tuning.ipynb](https://github.com/MohammedAly22/metro-asr/blob/main/examples/fine_tuning.ipynb) (Colab, needs a GPU), or see [Fine-tuning](https://github.com/MohammedAly22/metro-asr#fine-tuning) in the README. --- ## Limitations - **Dialect.** Trained on Egyptian Arabic; Modern Standard Arabic and other dialects — Gulf, Levantine, Maghrebi — degrade, Maghrebi most of all. - **Code-switching is Arabic-English only,** and it's the weakest part of the system: - Technical/domain vocabulary degrades under plain greedy decoding — see [The language head](#the-language-head) above for the fix and the measured improvement. - A lone English word surrounded by Arabic is harder than a full English clause, which tends to survive intact. - Acronyms and short initialisms (`CNN`, `MRI`, `AIC`) are acoustically ambiguous and unreliable without a matching language head. - Roughly 12K code-switching training utterances were available against ~130K Arabic-only ones — the imbalance is a data problem, not an architectural one. - **Clip length.** Trained and evaluated on 0.5–30 s; segment longer recordings before transcribing. - **Output is lowercase and unpunctuated**, matching the training transcripts. - **Not a streaming model in the strict sense** — the encoder is bidirectional, so a chunk must be complete before it can be decoded. `transcribe_stream` is chunked offline decoding, with a floor latency of one chunk. - **The language head is large.** The shipped 5-gram is 5.9 GB resident in RAM; a smaller 4-gram (like the domain heads above) trades some accuracy for a much smaller footprint. - **The language head asserts priors.** It corrects toward what its training text considers likely — which is exactly what makes it useful, and exactly how it gets unfamiliar proper nouns wrong. Train a head on your own text if this matters for your use case. - **Medium and Large are unreleased.** Their configs exist; no weights do. --- ## Citation ```bibtex @software{metro_asr_2025, title = {Metro-ASR: Non-Autoregressive Speech Recognition for Egyptian Arabic and Code-Switching with a Detachable N-gram Language Head}, author = {Mohammed Aly}, year = {2025}, url = {https://github.com/MohammedAly22/metro-asr} } ``` ## License MIT — see [LICENSE](https://github.com/MohammedAly22/metro-asr/blob/main/LICENSE).