nepali-conformer-streaming

Cache-aware streaming Nepali ASR (520 ms lookahead). Carries a large, honestly-reported streaming-lineage penalty on real calls — read RESULTS.md before choosing this over the offline model; it exists because a phone agent needs incremental output.

Try it: demo Space · Everything else: github.com/Ampixa/nepaliconformer (NepTel benchmark, per-system outputs, full honest results)

Numbers (measured, not marketed)

benchmark WER
NepTel — real Nepali call audio, human-reviewed refs 59.87
Held-out gold read Nepali (W1 slice) 31.5
Whisper-large-v3 zero-shot on the same NepTel audio 96.3

Architecture

121.3M-parameter 17-layer Conformer (d=512, striding ×4, 40 ms frames), hybrid TDT/CTC decoder, 1,024-piece Devanagari SentencePiece. Chunked-limited attention [[70,13],[70,6],[70,1],[70,0]], fully causal convolutions, cache-aware incremental decoding.

Training data

~1,655 h of mostly conversational Nepali (YouTube podcasts/interviews) with Google Chirp 2 pseudo-labels + 105 h human-labeled read speech; telephony codec, noise, reverb and tempo augmentation. Label-noise ceiling and every measured limitation (English, sung speech, slow speech, end-of-turn) are documented in the repo's RESULTS.md.

Usage

from nemo.collections.asr.models import EncDecHybridRNNTCTCBPEModel
m = EncDecHybridRNNTCTCBPEModel.restore_from("nepali_conformer_streaming.nemo")
print(m.transcribe(["audio.wav"])[0].text)

License: CC-BY-NC-4.0 (weights). Code in the repo: MIT.

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Real-call WER on NepTel v0.1 (real Nepali call-center audio, human-reviewed)
    self-reported
    59.870
  • Real-call CER on NepTel v0.1 (real Nepali call-center audio, human-reviewed)
    self-reported
    41.080
  • Read-speech WER on Held-out gold read Nepali (W1 read slice, OpenSLR-54 utterances absent from training)
    self-reported
    31.500