Automatic Speech Recognition
PyTorch
Safetensors
conformer
ctc
streaming
ghana
bfloat16

Griot Edge

Griot Edge is a 60.9M-parameter streaming Conformer CTC model for Akan, Dagbani, Dagaare, Ewe, Fante, Ghanaian English, and Ga. Its 151-character tokenizer preserves case, punctuation, and orthographic vowels.

Results

Scores normalize references and predictions with NFC, lowercase, and tone mark removal. Greedy uses CTC argmax; beam uses width 25 and the fused multilingual 3-gram KenLM with alpha 0.5 and beta 1.0.

Evaluation Decoding CER WER
Internal test, 109,008 utterances, causal Greedy 15.80% 37.39%
Internal test, 109,008 utterances, causal Beam + KenLM 15.81% 28.86%
Internal test, 109,008 utterances, 325 ms Greedy 15.04% 34.94%
Internal test, 109,008 utterances, 325 ms Beam + KenLM 14.59% 26.93%

Use

uvx --from huggingface-hub hf download Qlerqly/griot-edge --local-dir griot-edge
cd griot-edge
uv run --with-requirements requirements.txt python inference.py \
  --model-dir . --audio recording.wav --lookahead-frames 13

--lookahead-frames 0 selects causal inference. The evaluated 13-frame tier provides 325 ms of acoustic lookahead. The 6- and 26-frame tiers are available; 51 frames and full context are outside the validated release tiers.

The matching fused KenLM binary and its unigram list are included under kenlm/. File inference uses beam search with this LM by default (beam width 25, alpha 0.5, beta 1.0). Pass --greedy to disable it, or --kenlm PATH to select another language model. The LM guides decoding; it does not alter the acoustic weights.

Batch inference

The batch engine accepts multiple recordings and splits long inputs into 30-second acoustic windows with 5 seconds of overlap. Chunks from all inputs share batches, so one long recording can be processed in parallel on the GPU. Overlapping CTC frames are trimmed and joined before each recording is decoded; results retain the original input order. KenLM beam decoding remains the default.

After downloading the repository, create an inference environment:

cd griot-edge
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements-server.txt

Run a batch once, without starting a server:

.venv/bin/python batch_inference.py --model-dir . \
  --audio recording-a.wav recording-b.flac --batch-size 8 --output results.json

For Apple GPU inference add --device mps --dtype bf16; for NVIDIA GPU inference add --device cuda --dtype bf16. Use --device cpu --dtype fp32 for CPU inference. --greedy disables KenLM. A JSON manifest and reusable Python BatchEngine API are also available.

An optional server keeps one model and LM loaded and batches chunks across queued requests:

.venv/bin/python batch_server.py --model-dir . --device mps --dtype bf16 \
  --batch-size 8 --cpu-threads 6 --port 8000

It binds to loopback by default. GET /health reports readiness; POST /transcribe accepts {"inputs": [{"id": "clip", "audio_base64": "..."}]} with base64 WAV/FLAC file bytes. Replace the device/dtype flags as above for CPU or CUDA. This server is optional; the CLI and library need no server.

Batch sizes 1–128 were tested on Apple M5 Pro GPU and Modal NVIDIA L4 in BF16, and 1–8 on Apple M5 Pro CPU in FP32. Start at batch 8; larger batches did not improve throughput in the tested large workloads. These are synthetic speech capacity tests, not accuracy scores. CUDA BF16 showed batch-dependent transcripts on L4; evaluate representative speech accuracy before a CUDA release. Explicit BF16 never silently falls back to FP32.

See batch usage and HTTP client examples, native/container deployment and request limits, and the benchmark report. HTTP request limits must be raised separately when submitting 128 files in one request.

Server batching wait

The worker collects requests for up to 20 ms after taking the first queued request, or starts sooner once it has 8 requests in the group. The timer does not reset when another request arrives. It runs whatever it has collected when the deadline expires; it does not require a full acoustic batch. This wait applies once per request group, not once per 30-second chunk or acoustic batch. A single request already containing a full batch still uses the collection window unless it is disabled.

Set --batch-wait-ms 0 for immediate processing when the worker is available; --batch-wait-ms 20 is the default. --max-group-requests changes the request-group limit independently of --batch-size. Requests arriving during inference queue behind the current group, so their total queue time can exceed 20 ms. The separate --request-timeout defaults to 300 seconds.

Batch-size sweep results

Measured median wall-clock seconds, one warmup and three timed runs per configuration. Preprocessing, acoustic inference, score transfer, and decoding are included; model loading, file reading, and HTTP transfer are excluded. These are repeated synthetic English speech capacity tests, not WER results. Apple M5 Pro has 48 GiB unified memory; CPU and MPS runs use six PyTorch threads. Modal L4 runs use four allocated CPU cores. The different host CPUs affect end-to-end decoding times.

Full transcription with KenLM beam decoding

CPU uses FP32; both GPUs use BF16. CPU sweeps stop at batch 8.

Eight 30-second inputs — 240 seconds of audio

Batch limit Apple M5 Pro CPU FP32 (s) Apple M5 Pro GPU BF16 (s) Modal NVIDIA L4 BF16 (s)
1 5.149 0.741 2.633
2 5.272 0.697 2.480
4 4.592 0.660 2.342
8 3.145 0.675 2.317

Five mixed inputs (6, 12, 20, 30, 68 seconds) — 136 seconds, 7 chunks

Batch limit Apple M5 Pro CPU FP32 (s) Apple M5 Pro GPU BF16 (s) Modal NVIDIA L4 BF16 (s)
1 4.083 0.487 1.709
2 4.175 0.459 1.605
4 5.106 0.447 1.549
8 3.916 0.447 1.494

One four-minute recording — 240 seconds, 10 chunks

Batch limit Apple M5 Pro CPU FP32 (s) Apple M5 Pro GPU BF16 (s) Modal NVIDIA L4 BF16 (s)
1 6.766 1.352 4.944
2 6.999 1.295 4.748
4 6.884 1.287 4.676
8 5.409 1.257 4.634

Large GPU capacity sweep with greedy decoding

Both GPUs use BF16. Greedy decoding isolates GPU scaling from KenLM cost; these timings must not be treated as beam-decoding service throughput. Requested batch limits 1–128 are tested on sufficiently large workloads.

128 thirty-second inputs — 3,840 seconds of audio, 128 chunks

Batch limit Apple M5 Pro GPU BF16 (s) Modal NVIDIA L4 BF16 (s)
1 4.187 6.919
2 3.450 3.824
4 3.060 2.368
8 2.908 2.019
16 2.951 2.218
32 2.975 2.398
64 3.199 2.606
128 3.447 2.661

100 mixed inputs — 2,720 seconds of audio, 140 chunks

Batch limit Apple M5 Pro GPU BF16 (s) Modal NVIDIA L4 BF16 (s)
1 3.800 7.656
2 2.941 4.023
4 2.468 2.217
8 2.344 1.672
16 2.552 1.688
32 2.644 1.841
64 3.053 2.026
128 3.399 2.268

One hour-long recording — 3,600 seconds of audio, 144 chunks

Batch limit Apple M5 Pro GPU BF16 (s) Modal NVIDIA L4 BF16 (s)
1 5.103 7.839
2 4.303 4.274
4 3.873 2.496
8 3.781 2.093
16 3.782 2.294
32 3.830 2.585
64 4.137 2.777
128 3.933 2.843

All 24 large-workload configurations per GPU completed without OOM. Actual batches of 128 were exercised on the uniform and hour-long workloads; duration bucketing limited the mixed workload to at most 100 rows in a batch. M5 Pro transcripts matched batch-size-1 results throughout; L4 BF16 transcripts were batch dependent, as detailed in the numerical diagnostic. Start at batch 8.

For 128 × 30-second inputs with full KenLM beam decoding at batch 128, M5 Pro GPU took 10.423 s and Modal L4 took 38.318 s. On L4, decoding accounted for about 35.639 s, compared with 2.049 s for acoustic inference. This is why faster GPU acoustics did not translate into faster full transcripts.

See the benchmark report and raw JSON for timings, memory definitions, numerical checks, and limitations.

Flashlight decoding and optimized container

All inference paths now support --decoder-backend pyctcdecode or --decoder-backend flashlight; --greedy disables the LM. The Docker image includes both backends and starts with the fast tested preset:

Container setting Default
Decoder Flashlight with token beam 8
Acoustic batch limit 8
Decoder beam width / threshold 25 / 10
Decoder workers 1
Device and precision Auto: CPU FP32, supported GPU BF16
Server collection window 20 ms, up to 8 requests

Native CLI, library, and server commands retain pyctcdecode as their compatibility default. Select Flashlight explicitly there; its token beam defaults to 8. Flashlight constrains unknown words to its lexicon and changes transcripts. Its speed benefit is not an accuracy guarantee. Native MPS inference runs on macOS; Docker does not expose the Apple GPU.

Install the optional native decoder for Python usage:

uv pip install --python .venv/bin/python -r requirements-flashlight.txt
.venv/bin/python batch_inference.py --model-dir . --device mps --dtype bf16 \
  --audio recording-a.wav recording-b.flac --decoder-backend flashlight

The original inference.py file/live command and batch_server.py accept the same decoder selection. Native token beam 16 and 151, beam width, LM weights, custom LM/unigrams, and backend-specific pruning controls are exposed as flags.

Build and run the CPU container with all decoder options installed:

docker build -t griot-edge-batch:cpu .
docker run --rm --mount "type=bind,src=$PWD,dst=/model,readonly" \
  -p 127.0.0.1:8000:8000 --stop-timeout 300 griot-edge-batch:cpu

For NVIDIA GPU inference, build CUDA wheels on the Linux GPU host and require BF16:

docker build --build-arg TORCH_INDEX_URL=https://download.pytorch.org/whl/cu130 \
  -t griot-edge-batch:cuda .
docker run --rm --gpus all --mount "type=bind,src=$PWD,dst=/model,readonly" \
  -p 127.0.0.1:8000:8000 --stop-timeout 300 griot-edge-batch:cuda \
  server --device cuda --dtype bf16 --cpu-threads 4

Use cli --audio /mounted/audio.wav instead of server for one-shot work. Environment variables and command-line overrides support all engine/server controls; flags override environment values. For the reference decoder, set DECODER_BACKEND=pyctcdecode; for greedy set GREEDY=1. See deployment modes and full configuration table.

New matched Flashlight benchmarks

128 × 30-second inputs (64 minutes of audio), actual batch 128. Median of three runs after one warmup; preprocessing, acoustic inference, score transfer, and decoding included. Model loading and file/HTTP I/O excluded. GPU acoustics use BF16; both CPU decoders receive FP32 scores. Beam width 25; native token beam 8. These are synthetic English capacity measurements, not WER scores.

Hardware pyctcdecode total (s) Flashlight total (s) Speedup
Apple M5 Pro GPU + M5 Pro CPU, 6 Torch threads 10.524 5.056 2.08×
NVIDIA L4 + Modal host CPU, 4 allocated cores 26.127 13.321 1.96×

Matched end-to-end decode stage times:

Hardware pyctcdecode (s) Flashlight token beam 8 (s)
Apple M5 Pro CPU, BF16 acoustics on M5 Pro GPU 7.374 1.903
Modal host CPU, BF16 acoustics on L4 23.652 10.903

Decoder-only comparison on identical cached acoustic scores, separately measured for 128 × 30-second inputs:

Host CPU pyctcdecode (s) Native token beam 151 (s) Token beam 16 (s) Token beam 8 (s)
Apple M5 Pro CPU 7.304 11.515 3.492 1.936
Modal host CPU (4 cores) 22.747 79.442 13.494 6.264

Native search without token pruning was slower on these uniform workloads. The useful preset combines native search with token pruning. Transcripts differ from pyctcdecode, and names/unseen words need representative multilingual accuracy testing. Ten pathological lexicon entries longer than 128 characters are excluded from native trie construction; the binary LM is unchanged.

The new matched timings are a separate experiment from the earlier sweep results above. Compare backends within the same experiment: host CPU performance and CUDA BF16 acoustic scores can vary across runs and batch configurations. See Flashlight report and raw results for all workloads, search behavior, and validation.

Training and provenance

The model was pretrained for 44,000 audio hours, post-trained on no-speech/speech mixture for 12,000 hours, then fine-tuned with intermediate CTC and variable lookahead for 10,000 hours. See ATTRIBUTIONS.md and license_audit.json for source attribution. The source list follows Griot Nano 1 and adds KasaSpeech English–Twi code-switching speech and the University of Ghana Dagbani and Dagaare ASR subsets of WAXAL.

Limitations

The model is evaluated mainly on Ghanaian speech. Some Akan training transcripts were generated or corrected automatically, and performance on that subset is substantially worse than on human-transcribed Akan. Extreme lookahead settings were undertrained; use the documented tiers.

Downloads last month
71
Safetensors
Model size
60.9M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Qlerqly/griot-edge