PP-OCRv6 (tiny) multilingual text recognition base model

Hugging Face Hub mirror

This repository mirrors the canonical Zenodo release by Benjamin Kiessling, ALMAnaCH, Inria Paris. The weights are byte-identical to the Zenodo release and have not been modified by Small Models for GLAM.

Use from the Hugging Face Hub

Install Kraken and the Hugging Face Hub CLI:

pip install "kraken>=7.1.0" huggingface_hub

Download the model and run recognition:

hf download small-models-for-glam/kraken-ppocrv6-tiny \
    tiny.safetensors \
    --local-dir ./kraken-ppocrv6-tiny

kraken -i image.png output.txt \
    segment -bl \
    ocr -m ./kraken-ppocrv6-tiny/tiny.safetensors

Description

This is the tiny variant (~0.69M parameters) of a family of PP-OCRv6 text-line recognition models (tiny, small, medium) for kraken. The models are trained from scratch with baseline

  • bounding polygon data with a very diverse corpus containing historical, contemporary and born-digital document line images, handwritten and machine-printed, covering 44 languages across 10 scripts (Arabic, Armenian, Cyrillic, Ethiopic, Georgian, Greek, Hebrew, Latin, Malayalam, and Syriac).

Architecture

PP-OCRv6 is a conventional CTC-based line recognizer consisting of a lightweight convolutional backbone and a non-recurrent sequence-modelling neck.

The original architecture published as part of PaddlePaddle has been adapted for historical ATR by:

  • increasing line height to 128px
  • uncapping line width/CTC label budget
  • replacing the optimizer with Adam+Muon

Uses

This is a base model that is supposed to produce usable output across a wide range of scripts and materials out of the box while also allowing fine-tuning with ease. It should offer similar accuracy and generalization to VLM-based recognizers without hallucinations and with vastly higher throughput. This medium variant model achieves the highest scores on the test set, tiny and small trade accuracy for inference speed and a smaller memory footprint.

Transcription guidelines, Normalization, and Transformations

No attempt has been made to normalize the source datasets to a single set of transcription guidelines; the corpus mixes conventions, so inconsistent output is to be expected, in particular for Latin-script European manuscripts which mix large datasets such as CATMuS and TRIDIS that have very different approaches to transcription. Text was normalized to Unicode NFD and whitespace was normalized during training and evaluation.

Bias, Risks, and Limitations

The training corpus is heavily skewed towards a handful of high-resource languages (English, French, German, Latin, Dutch, Middle French, ...). Languages with little real training material show markedly higher error rates and will require fine-tuning for practical use. Because transcription conventions are inconsistent across sources, the model may resolve abbreviations or expand glyphs unpredictably.

The synthetic data used for training was created with the pangoline tool which is limited to approximating modern, machine-printed text. For the languages/scripts present only as synthetic data (Classical Armenian, Geʽez) and to a lesser extent those sharing the Latin script (Irish, Latvian, Lithuanian, Romanian, Serbian, Slovenian), real-world accuracy is probably limited.

How to Get Started with the Model

Install kraken (>= 7.1.0), download the model, and run recognition on an input image:

kraken -i image.png output.txt \
    segment -bl \
    ocr -m tiny.safetensors

For more information, refer to the documentation.

Training Details

Training Data

The model was trained on publicly available and restricted (private) datasets. Datasets marked as private are part of the training mixture but are not redistributable. A marks languages that were additionally augmented with synthetic training data (see below).

Language Script Datasets
Ancient Greek † Greek EPARCHOS, HPGTR, HTR_CPgr23, Stavronikita Monastery Greek Handwritten Document Collection No. 114, Stavronikita Monastery Greek Handwritten Document Collection No. 53, Stavronikita Monastery Greek Handwritten Document Collection No. 79, 11 private datasets
Arabic † Arabic Agapet, arabic_ms_data, iskandar, Muharaf: Manuscripts of Handwritten Arabic Dataset, OpenITI Arabic Print Data, RASAM dataset, TariMa
Catalan Latin FONDUE-CA-PRINT-20, htromance-spa, 1 private dataset
Church Slavonic Cyrillic 3 private datasets
Classical Armenian † Armenian synthetic only
Corsican Latin OCR Corse
Czech † Latin 2024--medieval-czech-main, 2023--medieval-czech, HTR Winter School 2025 - Medieval Czech - Biblioteka Jagiellonska BJ Rkp 441 IV, Paderov Bible handwriting ground truth, ehri
Danish Latin ehri
Dutch † Latin 6000 ground truth of VOC and notarial deeds / HTR of VOC, WIC and notarial deeds, ARletta, Dagboek Ernest Clarysse, FONDUE-NE-MSS-17-PR
English Latin FONDUE-EN-PRINT-20, IAM Handwriting Database, jcrs_train, jcrs_val, JosephHookerHTR, OCR-D gt_structure_text, sloanelab, The Revolutionary City / RevCity documentation, Memorials for Jane Lathrop Stanford, ehri, 2 private datasets
Finnish † Latin NewsEye/READ OCR Finnish Newspapers
French Latin Antoine Verard extracts, Copiste-d-un-jour, corpus-HTR-lignes-mixtes, dataset-celestine-doniau-danest, FONDUE-FR-MSS-19, FONDUE-FR-MSS-19-PR, FONDUE-FR-PRINT-20, FONDUE-MLT-ART, genauto-td-htr, HTR Front Justice, La Correspondance Doucet-Rene Jean, Memoire sur St Domingue par H. M. Michel, Moonshines, NewsEye READ AS French Newspapers, NuBIS-OCR, Recensement Valaisan (Valais Time Machine), Tapus Corpus, TIMEUS Corpus, TitresNobiliaires_17_18, CREMMA Manuscrits du 20e, CREMMA Wikipedia, Maxime Kovalewsky - Coutume contemporaine et loi ancienne (1893), HTRomance, Modern Roman languages corpus, Peraire Ground Truth, PARES, 1 private dataset
Georgian † Georgian 15 private datasets
German Latin 2024--medieval-german, 2025--Early-Modern-German, Bullinger Digital Gwalther handwriting ground truth, charlottenburger-amtsschrifttum, Chronicling Germany, dach-gt, DigiTheo Ground Truth, Dresdner Hofdiarium, Fibeln, FONDUE-DE-MSS-16-PR, FONDUE-DE-MSS-18, FONDUE-DE-MSS-19-PR, FONDUE-DE-MSS-20-PR, FONDUE-MLT-ART, FONDUE-MLT-PRINT-TEST, FoNDUE_Kunsthistorisches-UZH_Archivdatenbank, Ground truth for Neue Zurcher Zeitung black letter, Hakenkreuzbanner, inzigkofen, Klosterneuburg, Stiftsbibl., Cod. 48, koenigsfelden, mkn-kurrent-gt, NewsEye / READ OCR Austrian Newspapers, nuremberg_letterbooks, OCR-D gt_structure_text, reichsanzeiger-gt, Training Data Incunabula Reichenau, Weisthuemer, gt-fraktur, german_kurrent_handwritten_text_lines, Ground Truth (Tagebücher Edwin Hennig), ehri, Fanny loves Wilhelm, Frauen im Fokus, Graphemic Early New German, 1 private dataset
German (shorthand) Latin 1 private dataset
Geʽez † Ethiopic synthetic only
Hebrew Hebrew 2025-hebrew, 2 private datasets
Hungarian Latin ehri
Irish † Latin synthetic only
Italian Latin Diario del Soldato Bruno Celestino, EpiSearch HTR, FONDUE-IT-PRINT-20, FONDUE-IT-PRINT-20-PR, HTRogène Medieval Italian Manuscripts, HTRomance, Medieval Italian corpus of ground-truth for Handwritten Text Recognition, leopardi, LiDi1.0-project, LAM, 1 private dataset
Latin Latin 2025--late-medieval-latin-main, burchards-dekret-digital, Caroline Minuscule ground truth, Carolingian Latin Group HTR Wien Winter School 2025, CREMMA Medii Aevi, DISTINGUO Latin ground truth, Eutyches, FONDUE-LA-MSS-16-PR, FONDUE-LA-MSS-17-PR, FONDUE-LA-MSS-MA, FONDUE-LA-PRINT-16, HTR Winter School 2023/2024 - Late Medieval Latin, ONB 3891, HTR Winter School 2024/2025 - Late Medieval Latin, ONB 4135; ONB 4680, HTRogène Medieval Latin Manuscripts, HTRomance, Medieval Latin corpus of ground-truth for Handwritten Text Recognition, notarial_charter, nubis, OCR-D gt_structure_text, Paris Bible Project, Training Data Incunabula Reichenau, Wien ONB Cod 2160 ground truth
Latvian † Latin synthetic only
Lithuanian † Latin synthetic only
Malayalam Malayalam Ground Truth data for printed Malayalam
Middle Dutch Latin data
Middle French Latin Cremma Medieval, De la généalogie des dieux, Données imprimés du 16e siècle, Données HTR incunables du 15e siècle, Données HTR manuscrits du 15e siècle, Données imprimés du 18e siècle, Données imprimés gothiques du 16e siècle, Fabliaux, FONDUE-FR-AAEB-16, FONDUE-FR-AAEB-17, FONDUE-FR-MSS-18, FONDUE-FR-PRINT-16, FONDUE-FR-PRINT-17, HTR-SETAF-Jean-Michel, HTR-SETAF-LesFaictzJCH, HTR-SETAF-Pierre-de-Vingle, HTRogene French, HTRomance, Medieval French corpus of ground-truth for Handwritten Text Recognition, Imprimés 17e siècle, Liber, OCR17plus, TNAH-2021-DecameronFR, transcription-chastel
Multilingual (mixed) Latin Training Data Incunabula Reichenau, TranscriboQuest25_MedVernacReligio
Norwegian Latin NorHand v3 / Dataset for Handwritten Text Recognition in Norwegian
Occitan Latin HTRogène Medieval Occitan Manuscripts, 1 private dataset
Ottoman Turkish Arabic mehmed_ibn_mehmed_uskubi_cukrikcizade_altiparmak_risale, OpenITI Arabic-script OCR Catalyst Project print/typeface data
Persian † Arabic hafiz_divan, OpenITI Arabic-script OCR Catalyst Project print/typeface data, sadi_gulistan
Picard Latin 1 private dataset
Polish † Latin ehri
Portuguese Latin iForal-Dataset, Portuguese Handwriting 16th-19th c.
Romanian † Latin synthetic only
Russian † Cyrillic 2 private datasets
Serbian (Cyrillic) † Cyrillic synthetic only
Slovak Latin Slovensky Supermodel P&T1, ehri
Slovenian † Latin synthetic only
Spanish Latin FoNDUE Spanish chapbooks 19th c. Dataset, FONDUE-ES-MSS-19-PR, FONDUE-ES-PRINT-19, HTR - Araucania manuscript XIX, HTRogène Medieval Spanish Manuscripts, HTRomance, Medieval Spain corpus of ground-truth for Handwritten Text Recognition, ohg, 3 private datasets
Swedish Latin Finnish Court Records-sub500, kat57 Swedish ground truth dataset, NewsEye / READ OCR training dataset from Swedish Newspapers, riskarchiv
Syriac † Syriac zenodo.18157525, winter_school_vienna, 2 private datasets
Ukrainian Cyrillic 1 private dataset
Urdu Arabic OpenITI Arabic-script OCR Catalyst Project print/typeface data
Yiddish Hebrew 7 private datasets

Synthetic Training Data

Synthetic line images were generated as additional training material for 18 languages: Ancient Greek, Arabic, Classical Armenian, Czech, Dutch, Finnish, Georgian, Geʽez, Irish, Latvian, Lithuanian, Persian, Polish, Romanian, Russian, Serbian, Slovenian, Syriac. Of these, eight are present only as synthetic data: Classical Armenian, Geʽez, Irish, Latvian, Lithuanian, Romanian, Serbian (Cyrillic), Slovenian.

Training Procedure and Hyperparameters

The model was trained with kraken (feature/ppocrv6_rec branch).

Hardware 3 × NVIDIA H100
Precision bf16-mixed
Optimizer AdamW + Muon (momentum 0.95, weight decay 0.01)
Learning rate 1.7e-3 (cosine schedule, 1500 step warmup, min 1e-6)
Batch size 128
Gradient clipping 1.0
Epochs 16
Normalization NFD + whitespace
Augmentation enabled

Evaluation

Testing Data

Metrics are computed on a held-out test set for each language. The test split was obtained by random 5% split with an upper limit of 100 document pages per language. No attempt has been made to split in a manner that separates documents between train and test. The scores below are therefore best read as in-domain generalization.

Evaluations with purely synthetic data are marked with . CER and WER are the character- and word-level error rates, computed with torchmetrics using greedy CTC decoding and the NFD + whitespace normalization (equivalent to ketos test -u NFD -n).

Metrics

Language Lines CER (%) WER (%)
Ancient Greek 451 22.98 87.79
Arabic 1,926 20.11 68.77
Catalan 128 5.47 26.59
Church Slavonic 6,599 17.81 65.44
Classical Armenian ‡ 3,580 0.64 3.23
Corsican 50 1.88 13.33
Czech 922 15.47 61.28
Danish 25 1.29 10.89
Dutch 4,384 14.31 51.81
English 2,005 14.43 46.77
Finnish 7,907 1.16 6.83
French 5,967 14.02 31.40
Georgian 394 28.72 85.71
German 3,007 4.61 17.96
German (shorthand) 830 41.20 84.51
Geʽez ‡ 2,992 1.27 5.58
Hebrew 3,426 9.11 26.32
Hungarian 38 3.20 23.05
Irish ‡ 2,824 1.00 4.15
Italian 2,611 7.22 28.24
Latin 5,748 14.03 44.34
Latvian ‡ 2,397 1.25 6.48
Lithuanian ‡ 2,615 1.30 6.73
Malayalam 59 48.79 99.21
Middle Dutch 3,014 12.44 43.63
Middle French 3,970 8.70 32.38
Multilingual (mixed) 241 3.49 17.99
Norwegian 2,335 14.96 48.04
Ottoman Turkish 451 10.70 48.25
Persian 990 9.34 36.76
Polish 2,766 1.93 10.21
Portuguese 2,763 23.53 71.50
Romanian ‡ 2,412 1.50 6.87
Russian 3,053 25.11 68.33
Serbian (Cyrillic) ‡ 2,789 0.58 2.71
Slovak 1,550 5.11 21.50
Slovenian ‡ 2,741 2.42 6.87
Spanish 5,465 6.19 21.49
Swedish 3,555 11.89 45.04
Syriac 1,801 10.74 44.40
Ukrainian 1,253 13.71 47.56
Urdu 1,656 10.20 41.89
Yiddish 4,320 8.96 32.79
Aggregate (micro-average) 108,010 8.71 30.00
Aggregate (macro-average) 10.99 36.16

License

Released under the Apache-2.0 license.

Acknowledgements

Training of this model was funded by the European Union under Grant Agreement No.101132163 (ATRIUM) and No.101071829 (MiDRASH). Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.

This project also received funding from the BPI Scribe project.

Citation

If you use this model, please cite kraken and if possible credit the dataset providers linked in the front matter and the table above.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including small-models-for-glam/kraken-ppocrv6-tiny