- PP-OCRv6 (tiny) multilingual text recognition base model
PP-OCRv6 (tiny) multilingual text recognition base model
Hugging Face Hub mirror
This repository mirrors the canonical Zenodo release by Benjamin Kiessling, ALMAnaCH, Inria Paris. The weights are byte-identical to the Zenodo release and have not been modified by Small Models for GLAM.
Use from the Hugging Face Hub
Install Kraken and the Hugging Face Hub CLI:
pip install "kraken>=7.1.0" huggingface_hub
Download the model and run recognition:
hf download small-models-for-glam/kraken-ppocrv6-tiny \
tiny.safetensors \
--local-dir ./kraken-ppocrv6-tiny
kraken -i image.png output.txt \
segment -bl \
ocr -m ./kraken-ppocrv6-tiny/tiny.safetensors
Description
This is the tiny variant (~0.69M parameters) of a family of PP-OCRv6
text-line recognition models (tiny, small, medium) for
kraken. The models are trained from scratch with baseline
- bounding polygon data with a very diverse corpus containing historical, contemporary and born-digital document line images, handwritten and machine-printed, covering 44 languages across 10 scripts (Arabic, Armenian, Cyrillic, Ethiopic, Georgian, Greek, Hebrew, Latin, Malayalam, and Syriac).
Architecture
PP-OCRv6 is a conventional CTC-based line recognizer consisting of a lightweight convolutional backbone and a non-recurrent sequence-modelling neck.
The original architecture published as part of PaddlePaddle has been adapted for historical ATR by:
- increasing line height to 128px
- uncapping line width/CTC label budget
- replacing the optimizer with Adam+Muon
Uses
This is a base model that is supposed to produce usable output across a wide
range of scripts and materials out of the box while also allowing fine-tuning
with ease. It should offer similar accuracy and generalization to VLM-based
recognizers without hallucinations and with vastly higher throughput. This
medium variant model achieves the highest scores on the test set, tiny and
small trade accuracy for inference speed and a smaller memory footprint.
Transcription guidelines, Normalization, and Transformations
No attempt has been made to normalize the source datasets to a single set of transcription guidelines; the corpus mixes conventions, so inconsistent output is to be expected, in particular for Latin-script European manuscripts which mix large datasets such as CATMuS and TRIDIS that have very different approaches to transcription. Text was normalized to Unicode NFD and whitespace was normalized during training and evaluation.
Bias, Risks, and Limitations
The training corpus is heavily skewed towards a handful of high-resource languages (English, French, German, Latin, Dutch, Middle French, ...). Languages with little real training material show markedly higher error rates and will require fine-tuning for practical use. Because transcription conventions are inconsistent across sources, the model may resolve abbreviations or expand glyphs unpredictably.
The synthetic data used for training was created with the pangoline tool which is limited to approximating modern, machine-printed text. For the languages/scripts present only as synthetic data (Classical Armenian, Geʽez) and to a lesser extent those sharing the Latin script (Irish, Latvian, Lithuanian, Romanian, Serbian, Slovenian), real-world accuracy is probably limited.
How to Get Started with the Model
Install kraken (>= 7.1.0), download the model, and run recognition on an
input image:
kraken -i image.png output.txt \
segment -bl \
ocr -m tiny.safetensors
For more information, refer to the documentation.
Training Details
Training Data
The model was trained on publicly available and restricted (private) datasets. Datasets marked as private are part of the training mixture but are not redistributable. A † marks languages that were additionally augmented with synthetic training data (see below).
Synthetic Training Data
Synthetic line images were generated as additional training material for 18 languages: Ancient Greek, Arabic, Classical Armenian, Czech, Dutch, Finnish, Georgian, Geʽez, Irish, Latvian, Lithuanian, Persian, Polish, Romanian, Russian, Serbian, Slovenian, Syriac. Of these, eight are present only as synthetic data: Classical Armenian, Geʽez, Irish, Latvian, Lithuanian, Romanian, Serbian (Cyrillic), Slovenian.
Training Procedure and Hyperparameters
The model was trained with kraken (feature/ppocrv6_rec branch).
| Hardware | 3 × NVIDIA H100 |
| Precision | bf16-mixed |
| Optimizer | AdamW + Muon (momentum 0.95, weight decay 0.01) |
| Learning rate | 1.7e-3 (cosine schedule, 1500 step warmup, min 1e-6) |
| Batch size | 128 |
| Gradient clipping | 1.0 |
| Epochs | 16 |
| Normalization | NFD + whitespace |
| Augmentation | enabled |
Evaluation
Testing Data
Metrics are computed on a held-out test set for each language. The test split was obtained by random 5% split with an upper limit of 100 document pages per language. No attempt has been made to split in a manner that separates documents between train and test. The scores below are therefore best read as in-domain generalization.
Evaluations with purely synthetic data are marked with ‡. CER and WER are
the character- and word-level error rates, computed with torchmetrics using
greedy CTC decoding and the NFD + whitespace normalization (equivalent to
ketos test -u NFD -n).
Metrics
| Language | Lines | CER (%) | WER (%) |
|---|---|---|---|
| Ancient Greek | 451 | 22.98 | 87.79 |
| Arabic | 1,926 | 20.11 | 68.77 |
| Catalan | 128 | 5.47 | 26.59 |
| Church Slavonic | 6,599 | 17.81 | 65.44 |
| Classical Armenian ‡ | 3,580 | 0.64 | 3.23 |
| Corsican | 50 | 1.88 | 13.33 |
| Czech | 922 | 15.47 | 61.28 |
| Danish | 25 | 1.29 | 10.89 |
| Dutch | 4,384 | 14.31 | 51.81 |
| English | 2,005 | 14.43 | 46.77 |
| Finnish | 7,907 | 1.16 | 6.83 |
| French | 5,967 | 14.02 | 31.40 |
| Georgian | 394 | 28.72 | 85.71 |
| German | 3,007 | 4.61 | 17.96 |
| German (shorthand) | 830 | 41.20 | 84.51 |
| Geʽez ‡ | 2,992 | 1.27 | 5.58 |
| Hebrew | 3,426 | 9.11 | 26.32 |
| Hungarian | 38 | 3.20 | 23.05 |
| Irish ‡ | 2,824 | 1.00 | 4.15 |
| Italian | 2,611 | 7.22 | 28.24 |
| Latin | 5,748 | 14.03 | 44.34 |
| Latvian ‡ | 2,397 | 1.25 | 6.48 |
| Lithuanian ‡ | 2,615 | 1.30 | 6.73 |
| Malayalam | 59 | 48.79 | 99.21 |
| Middle Dutch | 3,014 | 12.44 | 43.63 |
| Middle French | 3,970 | 8.70 | 32.38 |
| Multilingual (mixed) | 241 | 3.49 | 17.99 |
| Norwegian | 2,335 | 14.96 | 48.04 |
| Ottoman Turkish | 451 | 10.70 | 48.25 |
| Persian | 990 | 9.34 | 36.76 |
| Polish | 2,766 | 1.93 | 10.21 |
| Portuguese | 2,763 | 23.53 | 71.50 |
| Romanian ‡ | 2,412 | 1.50 | 6.87 |
| Russian | 3,053 | 25.11 | 68.33 |
| Serbian (Cyrillic) ‡ | 2,789 | 0.58 | 2.71 |
| Slovak | 1,550 | 5.11 | 21.50 |
| Slovenian ‡ | 2,741 | 2.42 | 6.87 |
| Spanish | 5,465 | 6.19 | 21.49 |
| Swedish | 3,555 | 11.89 | 45.04 |
| Syriac | 1,801 | 10.74 | 44.40 |
| Ukrainian | 1,253 | 13.71 | 47.56 |
| Urdu | 1,656 | 10.20 | 41.89 |
| Yiddish | 4,320 | 8.96 | 32.79 |
| Aggregate (micro-average) | 108,010 | 8.71 | 30.00 |
| Aggregate (macro-average) | 10.99 | 36.16 |
License
Released under the Apache-2.0 license.
Acknowledgements
Training of this model was funded by the European Union under Grant Agreement
No.101132163 (ATRIUM) and No.101071829 (MiDRASH). Views and opinions
expressed are those of the authors only and do not necessarily reflect those of
the European Union.
This project also received funding from the BPI Scribe project.
Citation
If you use this model, please cite kraken and if possible credit the dataset providers linked in the front matter and the table above.