Instructions to use Horizon-Labs/punctuation-restoration-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Horizon-Labs/punctuation-restoration-small with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="Horizon-Labs/punctuation-restoration-small")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("Horizon-Labs/punctuation-restoration-small") model = AutoModelForTokenClassification.from_pretrained("Horizon-Labs/punctuation-restoration-small", device_map="auto") - Transformers.js
How to use Horizon-Labs/punctuation-restoration-small with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('token-classification', 'Horizon-Labs/punctuation-restoration-small'); - Notebooks
- Google Colab
- Kaggle
Punctuation Restoration (small, 140M): 88 languages
Adds punctuation back to text that has none: speech-recognition (ASR) transcripts, subtitles, chat messages, voice notes,
OCR or scraped text. It predicts, after each word, one of . , ? : - or nothing, in 88 languages, with any
casing (including the all-lowercase output of many ASR systems). Built on
mmBERT-small, Apache-2.0, ONNX files for CPU and the browser (transformers.js)
included. Try it in the browser.
A larger, more accurate version is available as punctuation-restoration-base.
- Same labels as oliverguhr/fullstop-punctuation-multilang-large
and kredor/punctuate-all (
0 . , ? - :), so it is a drop-in for thedeepmultilingualpunctuationpackage. - Chinese and Japanese work too (no word segmentation needed) with the included
punctuate.py, which writes,。?:there, and।(Hindi, Bengali),، ؟(Arabic script),;for questions in Greek and። ፣in Amharic. - It restores punctuation only. It does not change casing or the words.
onnx/model_quantized.onnx(int8 embeddings, 268 MB) predicts the same label as fp32 for 99.96% of tokens on held-out text.
Usage
With the included helper (any language; long texts are processed in overlapping chunks):
from huggingface_hub import hf_hub_download
import importlib.util, sys
spec = importlib.util.spec_from_file_location("punctuate", hf_hub_download("Horizon-Labs/punctuation-restoration-small", "punctuate.py"))
punctuate = importlib.util.module_from_spec(spec); spec.loader.exec_module(punctuate)
p = punctuate.Punctuator("Horizon-Labs/punctuation-restoration-small")
print(p("hello how are you today i hope everything is fine"))
print(p("今天天气很好我们去公园散步吧你觉得怎么样"))
As a drop-in for deepmultilingualpunctuation (space-separated languages; that package needs transformers 4.x, because it
passes grouped_entities, which transformers 5 removed):
from deepmultilingualpunctuation import PunctuationModel
model = PunctuationModel(model="Horizon-Labs/punctuation-restoration-small")
print(model.restore_punctuation("my name is anna i live in berlin what about you"))
Plain transformers: the label of each word is the label of its last token ("0" = no mark).
from transformers import pipeline
tagger = pipeline("token-classification", model="Horizon-Labs/punctuation-restoration-small")
tagger("hello how are you today")
transformers.js:
import { pipeline } from "@huggingface/transformers";
const tagger = await pipeline("token-classification", "Horizon-Labs/punctuation-restoration-small", { dtype: "q8" });
console.log(await tagger("hello how are you today i hope everything is fine"));
Evaluation
Punctuation F1: micro-averaged over the five marks (a predicted mark counts only if it is the correct mark at the correct
position), averaged over languages. Every model gets the same input: the paragraph with . , ? : ; ! - removed, once with its
original casing and once lowercased (as most ASR output is). Same script for every model (code/). The test sets were used only
for evaluation:
- FLORES-200 devtest: professionally translated news, travel and wiki text; the consecutive sentences of one article form a paragraph (about 280 paragraphs per language). Columns: the 12 languages of punctuate-all, and our 88 training languages.
- TED talks (TED2020 via OPUS): transcripts and their translations, 5 consecutive sentences, 200 paragraphs in each of 25 languages. Spoken style.
- Europarl: European Parliament proceedings, 5 consecutive sentences, 200 paragraphs in each of 12 languages. punctuate-all and the fullstop models were trained on Europarl, so this column is in-domain for them; we never trained on it.
| model | licence | languages | FLORES, 12 EU languages (cased) | (lowercased) | FLORES, our 88 languages (cased) | (lowercased) | TED talks, 25 languages (cased) | (lowercased) | Europarl, 12 languages (cased) | (lowercased) |
|---|---|---|---|---|---|---|---|---|---|---|
| this model (140M) | Apache-2.0 | 88 | 0.919 | 0.881 | 0.841 | 0.795 | 0.772 | 0.688 | 0.875 | 0.828 |
| punctuation-restoration-base (307M) | Apache-2.0 | 88 | 0.923 | 0.892 | 0.847 | 0.804 | 0.784 | 0.707 | 0.882 | 0.839 |
| punctuate-all | MIT | 12 | 0.879 | 0.873 | 0.667 | 0.657 | 0.664 | 0.654 | 0.885 | 0.884 |
| fullstop-punctuation-multilang-large | MIT | 4 | 0.841 | 0.838 | 0.669 | 0.663 | 0.648 | 0.642 | 0.826 | 0.824 |
| fullstop-punctuation-multilingual-base | MIT | 6 | 0.846 | 0.829 | 0.668 | 0.652 | 0.657 | 0.636 | 0.826 | 0.822 |
| xlm-roberta_punctuation_fullstop_truecase | Apache-2.0 | 47 | — | 0.818 | — | 0.689 | — | 0.643 | — | 0.751 |
| punct_cap_seg_47_language | Apache-2.0 | 47 | — | 0.565 | — | 0.448 | — | 0.510 | — | 0.504 |
- The 1-800-BAD-CODE models expect lowercased input and also restore casing and split sentences; they are scored only on
lowercased input, through the
punctuatorspackage, by re-reading their output against the input words. Words their output changed beyond casing (about 2% of words for xlm-roberta_punctuation_fullstop_truecase, 11% for punct_cap_seg_47_language) could not be aligned and count as "no mark", which lowers their scores somewhat. - On Europarl, punctuate-all (trained on it) is ahead of this model, most clearly on lowercased input.
Per mark (this model; mean over languages, FLORES over our 88):
| mark | FLORES (cased) | FLORES (lowercased) | TED (cased) | TED (lowercased) |
|---|---|---|---|---|
. |
0.945 | 0.882 | 0.890 | 0.775 |
, |
0.722 | 0.692 | 0.654 | 0.596 |
? |
0.778 | 0.629 | 0.784 | 0.686 |
: |
0.445 | 0.413 | 0.533 | 0.449 |
- |
0.174 | 0.183 | 0.064 | 0.055 |
Per language (FLORES)
| language | this model (cased) | this model (lowercased) | punctuate-all (cased) | xlm-roberta_punctuation_fullstop_truecase (lowercased) |
|---|---|---|---|---|
| Afrikaans | 0.845 | 0.809 | 0.738 | 0.767 |
| Albanian | 0.894 | 0.836 | 0.812 | 0.769 |
| Amharic | 0.768 | 0.768 | 0.557 | 0.721 |
| Arabic | 0.734 | 0.733 | 0.623 | 0.652 |
| Armenian | 0.608 | 0.575 | 0.540 | 0.551 |
| Azerbaijani | 0.867 | 0.819 | 0.685 | 0.753 |
| Basque | 0.855 | 0.792 | 0.576 | 0.543 |
| Belarusian | 0.936 | 0.881 | 0.823 | 0.808 |
| Bengali | 0.747 | 0.747 | 0.634 | 0.626 |
| Bosnian | 0.912 | 0.863 | 0.685 | 0.805 |
| Bulgarian | 0.927 | 0.889 | 0.881 | 0.849 |
| Catalan | 0.889 | 0.847 | 0.825 | 0.751 |
| Central Kurdish | 0.616 | 0.618 | 0.175 | 0.147 |
| Chinese (Simplified) | 0.832 | 0.834 | 0.576 | 0.771 |
| Croatian | 0.899 | 0.862 | 0.688 | 0.797 |
| Czech | 0.950 | 0.917 | 0.909 | 0.839 |
| Danish | 0.902 | 0.887 | 0.776 | 0.751 |
| Dutch | 0.927 | 0.897 | 0.899 | 0.835 |
| English | 0.873 | 0.827 | 0.810 | 0.780 |
| Esperanto | 0.900 | 0.851 | 0.779 | 0.713 |
| Estonian | 0.941 | 0.904 | 0.866 | 0.846 |
| Faroese | 0.871 | 0.826 | 0.438 | 0.466 |
| Filipino | 0.857 | 0.778 | 0.684 | 0.658 |
| Finnish | 0.940 | 0.898 | 0.859 | 0.864 |
| French | 0.908 | 0.876 | 0.861 | 0.806 |
| Galician | 0.896 | 0.854 | 0.833 | 0.752 |
| Georgian | 0.781 | 0.780 | 0.615 | 0.652 |
| German | 0.945 | 0.930 | 0.935 | 0.899 |
| Greek | 0.869 | 0.806 | 0.800 | 0.750 |
| Gujarati | 0.751 | 0.751 | 0.682 | 0.690 |
| Hausa | 0.771 | 0.590 | 0.499 | 0.512 |
| Hebrew | 0.789 | 0.788 | 0.744 | 0.611 |
| Hindi | 0.793 | 0.792 | 0.704 | 0.708 |
| Hungarian | 0.924 | 0.865 | 0.799 | 0.818 |
| Icelandic | 0.883 | 0.827 | 0.765 | 0.780 |
| Indonesian | 0.882 | 0.813 | 0.759 | 0.768 |
| Irish | 0.821 | 0.715 | 0.622 | 0.584 |
| Italian | 0.889 | 0.840 | 0.845 | 0.774 |
| Japanese | 0.906 | 0.906 | 0.446 | 0.806 |
| Javanese | 0.799 | 0.668 | 0.527 | 0.536 |
| Kannada | 0.746 | 0.745 | 0.648 | 0.709 |
| Kazakh | 0.903 | 0.846 | 0.711 | 0.752 |
| Kinyarwanda | 0.738 | 0.568 | 0.215 | 0.571 |
| Korean | 0.874 | 0.873 | 0.758 | 0.858 |
| Kyrgyz | 0.909 | 0.854 | 0.643 | 0.792 |
| Latvian | 0.942 | 0.899 | 0.845 | 0.861 |
| Lithuanian | 0.922 | 0.867 | 0.790 | 0.816 |
| Luxembourgish | 0.856 | 0.828 | 0.247 | 0.266 |
| Macedonian | 0.890 | 0.834 | 0.749 | 0.777 |
| Malay | 0.850 | 0.770 | 0.727 | 0.707 |
| Malayalam | 0.696 | 0.696 | 0.576 | 0.658 |
| Maltese | 0.837 | 0.760 | 0.188 | 0.246 |
| Marathi | 0.758 | 0.758 | 0.644 | 0.706 |
| Mongolian | 0.892 | 0.849 | 0.661 | 0.750 |
| Nepali | 0.820 | 0.819 | 0.686 | 0.686 |
| Norwegian Bokmål | 0.875 | 0.849 | 0.832 | 0.785 |
| Odia | 0.619 | 0.620 | 0.662 | 0.717 |
| Pashto | 0.727 | 0.728 | 0.522 | 0.652 |
| Persian | 0.811 | 0.812 | 0.745 | 0.678 |
| Polish | 0.927 | 0.883 | 0.898 | 0.857 |
| Portuguese | 0.903 | 0.853 | 0.848 | 0.793 |
| Punjabi | 0.770 | 0.770 | 0.681 | 0.685 |
| Romanian | 0.901 | 0.868 | 0.848 | 0.793 |
| Russian | 0.952 | 0.910 | 0.831 | 0.850 |
| Serbian | 0.897 | 0.849 | 0.711 | 0.798 |
| Sindhi | 0.728 | 0.730 | 0.587 | 0.646 |
| Sinhala | 0.723 | 0.723 | 0.624 | 0.731 |
| Slovak | 0.947 | 0.911 | 0.909 | 0.848 |
| Slovenian | 0.953 | 0.913 | 0.922 | 0.764 |
| Somali | 0.744 | 0.592 | 0.485 | 0.583 |
| Spanish | 0.878 | 0.840 | 0.835 | 0.778 |
| Sundanese | 0.786 | 0.631 | 0.472 | 0.489 |
| Swahili | 0.821 | 0.715 | 0.660 | 0.671 |
| Swedish | 0.877 | 0.855 | 0.837 | 0.796 |
| Tajik | 0.919 | 0.882 | 0.172 | 0.220 |
| Tamil | 0.714 | 0.713 | 0.599 | 0.688 |
| Tatar | 0.873 | 0.808 | 0.209 | 0.293 |
| Telugu | 0.723 | 0.723 | 0.628 | 0.695 |
| Turkish | 0.836 | 0.805 | 0.678 | 0.753 |
| Ukrainian | 0.928 | 0.894 | 0.846 | 0.848 |
| Urdu | 0.769 | 0.769 | 0.667 | 0.701 |
| Uyghur | 0.765 | 0.765 | 0.644 | 0.757 |
| Uzbek | 0.894 | 0.840 | 0.646 | 0.733 |
| Vietnamese | 0.849 | 0.769 | 0.718 | 0.668 |
| Welsh | 0.844 | 0.756 | 0.669 | 0.648 |
| Xhosa | 0.798 | 0.637 | 0.380 | 0.387 |
| Yoruba | 0.718 | 0.456 | 0.215 | 0.258 |
| Zulu | 0.802 | 0.659 | 0.360 | 0.365 |
Training
- Data: 1,168,626 text windows (119M words) in 88 languages from FineWeb-2 and FineWeb (ODC-BY). The labels come from the text's own punctuation: snippets that end in a full stop, are mostly letters and are not lists; 1-4 snippets per window (up to 230 words), 30% cut mid-sentence. 70% of the inputs were lowercased.
- Normalisation:
!,;,…and script full stops (。 । ። ۔...) count as.;,、،as,;?؟(and;in Greek) as?; a free-standing dash as-. A.before a lowercase word (abbreviations) counts as no mark. - Model: mmBERT-small token classification, one label per word on its last token, 2 epochs, learning rate 5e-5. Checkpoint and number of epochs (2 over 1) chosen on held-out web text, not on the test sets.
- Code:
code/in this repository.
Limitations
- Only
. , ? : -are predicted. Exclamation marks, semicolons, quotes, brackets and Spanish inverted marks are not restored.:and especially-(a free-standing dash) are much less reliable than.,?(see the per-mark table). - Accuracy on lowercased text is lower than on cased text, because capital letters reveal sentence starts.
- Trained on written web text. Disfluent speech (fillers, false starts, repetitions) is harder, and punctuation conventions vary by writer. Thai, Lao, Khmer, Burmese and Tibetan are not supported.
- Low-resource languages and languages outside the 88 score lower (see the per-language table).
- Downloads last month
- 42
Model tree for Horizon-Labs/punctuation-restoration-small
Base model
jhu-clsp/mmBERT-small