Oído: Conformer-CTC Small, int8, for the ESP32-S3

Animated demo: the transcripts are Oído's output (sped up). Footage from a physical board is coming.

¡Oído! is Spanish kitchen slang for heard, got it.

This is open-vocabulary English speech recognition that runs entirely on an ESP32-S3 microcontroller (240 MHz dual-core, 8 MB PSRAM, 16 MB flash), with no cloud and no neural accelerator. The file nemo8.tnm is NVIDIA's stt_en_conformer_ctc_small (13 M parameters), quantized to int8 and packed for Lokutor's on-chip engine.

Code, firmware and tools: github.com/lokutor-ai/oido (GPLv3, commercial licenses available).

LibriSpeech WER (%) test-clean test-other
This model on the ESP32-S3 engine (int8, greedy) 3.70 8.23
Original fp32 model (PyTorch) 3.68 8.11
Espressif MultiNet7 on the same chip (ESP-SR benchmark) 8.5 21.3
Whisper tiny.en, fp32 on a laptop (500-utterance subsets) 6.3 15.9
+ on-chip language model (nemo_lm.tlm, beam search, int8 greedy → beam 4) 3.33 7.15

nemo_lm.tlm is a 1.3 M-parameter GRU language model over the same 1024 tokens, trained only on the public-domain book text of the LibriSpeech LM corpus and released under CC-BY-4.0. It is optional: without it the engine decodes greedily.

Status (9 October 2026). Measured on an ESP32-S3-WROOM-1-N16R8 board, one core at 240 MHz: 1.97× real time (1.91–2.02 over 27 clips), so a 4 s command shows its text about 8.7 s after it ends, including the 0.8 s end-of-speech wait. That is not real time yet: the two-core mode, which emulation had predicted would approach real time, gives wrong transcripts on silicon and is disabled. The WERs are computed with the host build of the firmware engine (same C code, same int8 arithmetic). The board's transcripts are identical to the instruction-exact emulator's on all 27 clips; the laptop's math library rounds the last bit differently and changes a word on some hard clips (22 of 27 identical; the word error rate on those 27 clips is 7.14 % for the board and for the host build). Details and raw numbers: github.com/lokutor-ai/oido.

Need flash space for your own application? The int4 version is 8.3 MB (4.61 / 9.98 WER) and leaves a 6 MB app partition free.

Need lower latency? The streaming variant shows text while you speak; on the board its final text appeared 4.1–11.1 s after clips of 3.6–10.6 s ended (about half the wait of utterance mode), at some cost in accuracy.

Español: lokutor-ai/oido-es-ctc-small-int8 (CC-BY-4.0).

Which Oído model?

One model per language; the models that stream also run in full-context (utterance) mode.

Model Language Modes Size License
oido-ctc-small-int8 English utterance only (best accuracy) 14.0 MB (+1.3 MB LM) CC-BY-4.0
oido-ctc-small-int4 English utterance only (smallest) 8.3 MB CC-BY-SA-4.0
oido-ctc-small-stream-int8 English streaming + full-context (low latency) 14.0 MB CC-BY-SA-4.0
oido-es-ctc-small-int8 Spanish streaming + full-context in one file 14.0 MB (+1.3 MB LM) CC-BY-4.0

Streaming-capable files stream automatically in the firmware's microphone mode and in live_demo.py.

Use

git clone https://github.com/lokutor-ai/oido && cd oido
esp32/host/tasr_cli models/nemo8.tnm recording.wav          # laptop, same engine as the chip (after `make` in esp32/host)
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm          # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone

Files

  • nemo8.tnm: int8 weights in the TNM1 format (14.0 MB), produced by train/export_nemo.py in the GitHub repository.
  • tokenizer.model: the original SentencePiece tokenizer (1024 BPE pieces).

License and attribution

Derived from NVIDIA's stt_en_conformer_ctc_small, licensed CC-BY-4.0. Changes: re-implemented in PyTorch, quantized to int8 and repacked. This model is also released under CC-BY-4.0. The engine and firmware are GPLv3, with commercial licenses from Lokutor.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lokutor-ai/oido-ctc-small-int8

Finetuned
(5)
this model

Dataset used to train lokutor-ai/oido-ctc-small-int8

Evaluation results