Griot Edge
Griot Edge is a 60.9M-parameter streaming Conformer CTC model for Akan, Dagbani, Dagaare, Ewe, Fante, Ghanaian English, and Ga. Its 151-character tokenizer preserves case, punctuation, and orthographic vowels.
Results
Scores normalize references and predictions with NFC, lowercase, and tone mark removal. Greedy uses CTC argmax; beam uses width 25 and the fused multilingual 3-gram KenLM with alpha 0.5 and beta 1.0.
| Evaluation | Decoding | CER | WER |
|---|---|---|---|
| Internal test, 109,008 utterances, causal | Greedy | 15.80% | 37.39% |
| Internal test, 109,008 utterances, causal | Beam + KenLM | 15.81% | 28.86% |
| Internal test, 109,008 utterances, 325 ms | Greedy | 15.04% | 34.94% |
| Internal test, 109,008 utterances, 325 ms | Beam + KenLM | 14.59% | 26.93% |
Use
uvx --from huggingface-hub hf download Qlerqly/griot-edge --local-dir griot-edge
cd griot-edge
uv run --with-requirements requirements.txt python inference.py \
--model-dir . --audio recording.wav --lookahead-frames 13
--lookahead-frames 0 selects causal inference. The evaluated 13-frame tier
provides 325 ms of acoustic lookahead. The 6- and 26-frame tiers are available;
51 frames and full context are outside the validated release tiers.
The matching fused KenLM binary and its unigram list are included under
kenlm/. File inference uses beam search with this LM by default (beam width
25, alpha 0.5, beta 1.0). Pass --greedy to disable it, or --kenlm PATH to
select another language model. The LM guides decoding; it does not alter the
acoustic weights.
Batch inference
The batch engine accepts multiple recordings and splits long inputs into 30-second acoustic windows with 5 seconds of overlap. Chunks from all inputs share batches, so one long recording can be processed in parallel on the GPU. Overlapping CTC frames are trimmed and joined before each recording is decoded; results retain the original input order. KenLM beam decoding remains the default.
After downloading the repository, create an inference environment:
cd griot-edge
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements-server.txt
Run a batch once, without starting a server:
.venv/bin/python batch_inference.py --model-dir . \
--audio recording-a.wav recording-b.flac --batch-size 8 --output results.json
For Apple GPU inference add --device mps --dtype bf16; for NVIDIA GPU
inference add --device cuda --dtype bf16. Use --device cpu --dtype fp32
for CPU inference. --greedy disables KenLM. A JSON manifest and reusable
Python BatchEngine API are also available.
An optional server keeps one model and LM loaded and batches chunks across queued requests:
.venv/bin/python batch_server.py --model-dir . --device mps --dtype bf16 \
--batch-size 8 --cpu-threads 6 --port 8000
It binds to loopback by default. GET /health reports readiness;
POST /transcribe accepts {"inputs": [{"id": "clip", "audio_base64": "..."}]}
with base64 WAV/FLAC file bytes. Replace the device/dtype flags as above for
CPU or CUDA. This server is optional; the CLI and library need no server.
Batch sizes 1–128 were tested on Apple M5 Pro GPU and Modal NVIDIA L4 in BF16, and 1–8 on Apple M5 Pro CPU in FP32. Start at batch 8; larger batches did not improve throughput in the tested large workloads. These are synthetic speech capacity tests, not accuracy scores. CUDA BF16 showed batch-dependent transcripts on L4; evaluate representative speech accuracy before a CUDA release. Explicit BF16 never silently falls back to FP32.
See batch usage and HTTP client examples, native/container deployment and request limits, and the benchmark report. HTTP request limits must be raised separately when submitting 128 files in one request.
Server batching wait
The worker collects requests for up to 20 ms after taking the first queued request, or starts sooner once it has 8 requests in the group. The timer does not reset when another request arrives. It runs whatever it has collected when the deadline expires; it does not require a full acoustic batch. This wait applies once per request group, not once per 30-second chunk or acoustic batch. A single request already containing a full batch still uses the collection window unless it is disabled.
Set --batch-wait-ms 0 for immediate processing when the worker is available;
--batch-wait-ms 20 is the default. --max-group-requests changes the
request-group limit independently of --batch-size. Requests arriving during
inference queue behind the current group, so their total queue time can exceed
20 ms. The separate --request-timeout defaults to 300 seconds.
Batch-size sweep results
Measured median wall-clock seconds, one warmup and three timed runs per configuration. Preprocessing, acoustic inference, score transfer, and decoding are included; model loading, file reading, and HTTP transfer are excluded. These are repeated synthetic English speech capacity tests, not WER results. Apple M5 Pro has 48 GiB unified memory; CPU and MPS runs use six PyTorch threads. Modal L4 runs use four allocated CPU cores. The different host CPUs affect end-to-end decoding times.
Full transcription with KenLM beam decoding
CPU uses FP32; both GPUs use BF16. CPU sweeps stop at batch 8.
Eight 30-second inputs — 240 seconds of audio
| Batch limit | Apple M5 Pro CPU FP32 (s) | Apple M5 Pro GPU BF16 (s) | Modal NVIDIA L4 BF16 (s) |
|---|---|---|---|
| 1 | 5.149 | 0.741 | 2.633 |
| 2 | 5.272 | 0.697 | 2.480 |
| 4 | 4.592 | 0.660 | 2.342 |
| 8 | 3.145 | 0.675 | 2.317 |
Five mixed inputs (6, 12, 20, 30, 68 seconds) — 136 seconds, 7 chunks
| Batch limit | Apple M5 Pro CPU FP32 (s) | Apple M5 Pro GPU BF16 (s) | Modal NVIDIA L4 BF16 (s) |
|---|---|---|---|
| 1 | 4.083 | 0.487 | 1.709 |
| 2 | 4.175 | 0.459 | 1.605 |
| 4 | 5.106 | 0.447 | 1.549 |
| 8 | 3.916 | 0.447 | 1.494 |
One four-minute recording — 240 seconds, 10 chunks
| Batch limit | Apple M5 Pro CPU FP32 (s) | Apple M5 Pro GPU BF16 (s) | Modal NVIDIA L4 BF16 (s) |
|---|---|---|---|
| 1 | 6.766 | 1.352 | 4.944 |
| 2 | 6.999 | 1.295 | 4.748 |
| 4 | 6.884 | 1.287 | 4.676 |
| 8 | 5.409 | 1.257 | 4.634 |
Large GPU capacity sweep with greedy decoding
Both GPUs use BF16. Greedy decoding isolates GPU scaling from KenLM cost; these timings must not be treated as beam-decoding service throughput. Requested batch limits 1–128 are tested on sufficiently large workloads.
128 thirty-second inputs — 3,840 seconds of audio, 128 chunks
| Batch limit | Apple M5 Pro GPU BF16 (s) | Modal NVIDIA L4 BF16 (s) |
|---|---|---|
| 1 | 4.187 | 6.919 |
| 2 | 3.450 | 3.824 |
| 4 | 3.060 | 2.368 |
| 8 | 2.908 | 2.019 |
| 16 | 2.951 | 2.218 |
| 32 | 2.975 | 2.398 |
| 64 | 3.199 | 2.606 |
| 128 | 3.447 | 2.661 |
100 mixed inputs — 2,720 seconds of audio, 140 chunks
| Batch limit | Apple M5 Pro GPU BF16 (s) | Modal NVIDIA L4 BF16 (s) |
|---|---|---|
| 1 | 3.800 | 7.656 |
| 2 | 2.941 | 4.023 |
| 4 | 2.468 | 2.217 |
| 8 | 2.344 | 1.672 |
| 16 | 2.552 | 1.688 |
| 32 | 2.644 | 1.841 |
| 64 | 3.053 | 2.026 |
| 128 | 3.399 | 2.268 |
One hour-long recording — 3,600 seconds of audio, 144 chunks
| Batch limit | Apple M5 Pro GPU BF16 (s) | Modal NVIDIA L4 BF16 (s) |
|---|---|---|
| 1 | 5.103 | 7.839 |
| 2 | 4.303 | 4.274 |
| 4 | 3.873 | 2.496 |
| 8 | 3.781 | 2.093 |
| 16 | 3.782 | 2.294 |
| 32 | 3.830 | 2.585 |
| 64 | 4.137 | 2.777 |
| 128 | 3.933 | 2.843 |
All 24 large-workload configurations per GPU completed without OOM. Actual batches of 128 were exercised on the uniform and hour-long workloads; duration bucketing limited the mixed workload to at most 100 rows in a batch. M5 Pro transcripts matched batch-size-1 results throughout; L4 BF16 transcripts were batch dependent, as detailed in the numerical diagnostic. Start at batch 8.
For 128 × 30-second inputs with full KenLM beam decoding at batch 128, M5 Pro GPU took 10.423 s and Modal L4 took 38.318 s. On L4, decoding accounted for about 35.639 s, compared with 2.049 s for acoustic inference. This is why faster GPU acoustics did not translate into faster full transcripts.
See the benchmark report and raw JSON for timings, memory definitions, numerical checks, and limitations.
Flashlight decoding and optimized container
All inference paths now support --decoder-backend pyctcdecode or
--decoder-backend flashlight; --greedy disables the LM. The Docker image
includes both backends and starts with the fast tested preset:
| Container setting | Default |
|---|---|
| Decoder | Flashlight with token beam 8 |
| Acoustic batch limit | 8 |
| Decoder beam width / threshold | 25 / 10 |
| Decoder workers | 1 |
| Device and precision | Auto: CPU FP32, supported GPU BF16 |
| Server collection window | 20 ms, up to 8 requests |
Native CLI, library, and server commands retain pyctcdecode as their compatibility default. Select Flashlight explicitly there; its token beam defaults to 8. Flashlight constrains unknown words to its lexicon and changes transcripts. Its speed benefit is not an accuracy guarantee. Native MPS inference runs on macOS; Docker does not expose the Apple GPU.
Install the optional native decoder for Python usage:
uv pip install --python .venv/bin/python -r requirements-flashlight.txt
.venv/bin/python batch_inference.py --model-dir . --device mps --dtype bf16 \
--audio recording-a.wav recording-b.flac --decoder-backend flashlight
The original inference.py file/live command and batch_server.py accept the
same decoder selection. Native token beam 16 and 151, beam width, LM weights,
custom LM/unigrams, and backend-specific pruning controls are exposed as flags.
Build and run the CPU container with all decoder options installed:
docker build -t griot-edge-batch:cpu .
docker run --rm --mount "type=bind,src=$PWD,dst=/model,readonly" \
-p 127.0.0.1:8000:8000 --stop-timeout 300 griot-edge-batch:cpu
For NVIDIA GPU inference, build CUDA wheels on the Linux GPU host and require BF16:
docker build --build-arg TORCH_INDEX_URL=https://download.pytorch.org/whl/cu130 \
-t griot-edge-batch:cuda .
docker run --rm --gpus all --mount "type=bind,src=$PWD,dst=/model,readonly" \
-p 127.0.0.1:8000:8000 --stop-timeout 300 griot-edge-batch:cuda \
server --device cuda --dtype bf16 --cpu-threads 4
Use cli --audio /mounted/audio.wav instead of server for one-shot work.
Environment variables and command-line overrides support all engine/server
controls; flags override environment values. For the reference decoder, set
DECODER_BACKEND=pyctcdecode; for greedy set GREEDY=1. See
deployment modes and full configuration table.
New matched Flashlight benchmarks
128 × 30-second inputs (64 minutes of audio), actual batch 128. Median of three runs after one warmup; preprocessing, acoustic inference, score transfer, and decoding included. Model loading and file/HTTP I/O excluded. GPU acoustics use BF16; both CPU decoders receive FP32 scores. Beam width 25; native token beam 8. These are synthetic English capacity measurements, not WER scores.
| Hardware | pyctcdecode total (s) | Flashlight total (s) | Speedup |
|---|---|---|---|
| Apple M5 Pro GPU + M5 Pro CPU, 6 Torch threads | 10.524 | 5.056 | 2.08× |
| NVIDIA L4 + Modal host CPU, 4 allocated cores | 26.127 | 13.321 | 1.96× |
Matched end-to-end decode stage times:
| Hardware | pyctcdecode (s) | Flashlight token beam 8 (s) |
|---|---|---|
| Apple M5 Pro CPU, BF16 acoustics on M5 Pro GPU | 7.374 | 1.903 |
| Modal host CPU, BF16 acoustics on L4 | 23.652 | 10.903 |
Decoder-only comparison on identical cached acoustic scores, separately measured for 128 × 30-second inputs:
| Host CPU | pyctcdecode (s) | Native token beam 151 (s) | Token beam 16 (s) | Token beam 8 (s) |
|---|---|---|---|---|
| Apple M5 Pro CPU | 7.304 | 11.515 | 3.492 | 1.936 |
| Modal host CPU (4 cores) | 22.747 | 79.442 | 13.494 | 6.264 |
Native search without token pruning was slower on these uniform workloads. The useful preset combines native search with token pruning. Transcripts differ from pyctcdecode, and names/unseen words need representative multilingual accuracy testing. Ten pathological lexicon entries longer than 128 characters are excluded from native trie construction; the binary LM is unchanged.
The new matched timings are a separate experiment from the earlier sweep results above. Compare backends within the same experiment: host CPU performance and CUDA BF16 acoustic scores can vary across runs and batch configurations. See Flashlight report and raw results for all workloads, search behavior, and validation.
Training and provenance
The model was pretrained for 44,000 audio hours, post-trained on no-speech/speech mixture for 12,000 hours, then fine-tuned with intermediate CTC and variable lookahead for 10,000 hours. See ATTRIBUTIONS.md and license_audit.json for source attribution. The source list follows Griot Nano 1 and adds KasaSpeech English–Twi code-switching speech and the University of Ghana Dagbani and Dagaare ASR subsets of WAXAL.
Limitations
The model is evaluated mainly on Ghanaian speech. Some Akan training transcripts were generated or corrected automatically, and performance on that subset is substantially worse than on human-transcribed Akan. Extreme lookahead settings were undertrained; use the documented tiers.
- Downloads last month
- 71