Peppercorn (Hạt Tiêu), 0.3B

Peppercorn is a compact vision-language model for Vietnamese and English document OCR, with 328M parameters.

It takes a document page as input and produces structured Markdown, including HTML tables and LaTeX mathematical expressions.

Its name comes from the Vietnamese saying *nhỏ nhưng có võ — small, but mighty: compact in size, yet capable.

Compared with original design, Peppercorn uses approximately 4× fewer image tokens per page (~900 vs. ~3,600) and achieves higher throughput under our tested serving configuration.

It performs particularly well on Vietnamese administrative documents and synthetic tables, while English OCR, mathematical expressions, and some real-world financial tables remain areas for improvement.

Highlights

  • 328M parameters for efficient document understanding.
  • Vietnamese and English OCR, covering both born-digital documents and scanned pages.
  • Structured output in Markdown, HTML tables, and LaTeX.
  • Efficient visual encoding, using approximately 880 image tokens per page on average.
  • High-throughput inference with vLLM.
  • OpenAI-compatible API integration through the provided serving interface.

Intended Use

Peppercorn is designed for document digitization and structured text extraction, including:

  • Vietnamese administrative documents.
  • Scanned documents and born-digital pages.
  • Documents containing tables and structured layouts.
  • Mathematical expressions and technical documents.
  • Batch document processing pipelines.

The model generates text from document images. Its output should be validated before being used in workflows where transcription errors could have significant consequences.

Architecture

Peppercorn uses a compact vision-language architecture designed specifically for efficient document OCR.

Component Description
Vision encoder Extracts visual features from document images
Visual processing Image tiling and feature aggregation for page-level understanding
Connector Converts visual features into representations suitable for the language model
Text decoder Generates structured text from visual input
Output formats Markdown, HTML tables, and LaTeX

Image Processing

The longest image edge is resized to a maximum of 2,048 pixels. Larger pages are processed using image tiles, together with a global view when applicable.

This allows the model to process full document pages while controlling the number of visual tokens.

The average input is approximately 880 image tokens per page, although the actual number depends on the document dimensions and tiling configuration.

Output Format

The model is intended to produce readable, structured text rather than unformatted character sequences.

For example, tables may be represented as HTML:

<table>
  <tr>
    <th>Item</th>
    <th>Value</th>
  </tr>
  <tr>
    <td>Example</td>
    <td>123</td>
  </tr>
</table>

Mathematical expressions may be represented using LaTeX:

\frac{a+b}{c}

The exact output structure depends on the document content and model generation behavior.

Evaluation

All CER values below are calculated on normalized text. Lower is better.

Evaluation uses greedy decoding with a maximum of 4,096 generated tokens per page unless otherwise specified.

Evaluation Set Pages v1 Mean CER v2 Mean CER v2 Median CER v2 Looped Pages
Vietnamese born-digital 447 0.038 0.046 0.004 9
Vietnamese born-digital, hard 129 0.139 0.140 0.032 13
Vietnamese scans 189 0.062 0.071 0.011 7
English pages 280 0.127 0.140 0.032 14
English crops 100 0.013 0.027 0.000 1
Vietnamese crops 200 0.018 0.022 0.000 2
Vietnamese mathematics 162 0.077 0.088 0.023 7
Synthetic Vietnamese tables 400 0.023 0.014 0.000 1
Real financial tables (HTML) 85 0.082 0.106 0.047 2
All evaluation sets 1,992 0.058 0.064 0.006 56 (2.8%)
Administrative documents 732 0.0224 0.0169 — —
Administrative pages containing tables 141 0.057 0.043 — —

Interpretation

Peppercorn improves on the administrative document evaluation set and synthetic table recognition.

Across the complete evaluation suite, v1 retains a lower mean CER. The results indicate that the newer model's efficiency improvements do not translate into uniform accuracy gains across every document category.

A page is classified as looped if generation does not terminate within the token limit or produces output exceeding three times the reference length.

With the loop guard enabled, the v2 evaluation CER decreases to 0.061.

On pages where neither model loops, v2 performs comparably to v1.

These results are specific to the evaluation datasets and inference settings described here. They should not be interpreted as a guarantee of accuracy on arbitrary documents.

Throughput

Throughput was measured on an RTX PRO 6000 (96 GB) using vLLM 0.29 and a workload of 800 mixed Vietnamese document pages, including born-digital pages, scans, and tables.

The serving setup used the provided example scripts, streaming inference, and 1,024 concurrent requests.

Configuration Throughput Output Tokens per Page
1 engine, 8 API servers, no loop guard 14.9 pages/s 687
1 engine, 8 API servers, loop guard enabled 15.1 pages/s 610
2 engines on one GPU, 4 API servers each, loop guard enabled 22.1 pages/s 602

A single engine does not fully utilize the GPU under this workload. Running two engine replicas on the same GPU increases aggregate throughput by approximately 1.4× in the tested configuration.

Streaming individual tokens introduces additional overhead. The recommended serving configuration uses 32-token streaming intervals to balance responsiveness and throughput.

Actual performance varies with image dimensions, document complexity, output length, concurrency, GPU configuration, and serving parameters.

Installation

Requirements

  • Python
  • PyTorch
  • transformers>=5 for Transformers inference
  • vllm>=0.29 for the provided vLLM integration

Download the Model

pip install "vllm>=0.29"

hf download DuyTa/peppercorn-v2-0.3b \
    --local-dir peppercorn-v2

Install the provided vLLM integration:

pip install ./peppercorn-v2/vllm_plugin

The integration registers the model architecture with vLLM through its plugin mechanism.

Inference with Transformers

The following example loads the model and extracts text from a document image.

from transformers import AutoModelForImageTextToText, AutoProcessor
from PIL import Image
import torch

repo = "DuyTa/peppercorn-v2-0.3b"

processor = AutoProcessor.from_pretrained(repo)

model = AutoModelForImageTextToText.from_pretrained(
    repo,
    dtype=torch.bfloat16,
).cuda().eval()

image = Image.open("page.png").convert("RGB")

prompt = processor.apply_chat_template(
    [
        {
            "role": "user",
            "content": [{"type": "image"}],
        }
    ],
    add_generation_prompt=True,
)

inputs = processor(
    text=prompt,
    images=[[image]],
    return_tensors="pt",
).to("cuda")

inputs["pixel_values"] = inputs["pixel_values"].to(
    torch.bfloat16
)

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=4096,
        do_sample=False,
    )

generated_tokens = output[
    0,
    inputs["input_ids"].shape[1]:,
]

print(
    processor.decode(
        generated_tokens,
        skip_special_tokens=True,
    )
)

The prompt contains the document image without additional textual instructions.

A pad_token_id warning referencing 128002 may appear because of the default configuration inherited from the Idefics3 configuration. In the described setup, this warning is harmless.

Inference with vLLM

The provided vLLM integration supports offline inference and OpenAI-compatible serving.

Offline Inference

python peppercorn-v2/examples/offline_infer.py \
    --model DuyTa/peppercorn-v2-0.3b \
    page1.png page2.jpg

The script prints the extracted Markdown along with the generation finish reason.

Start the API Server

Run one replica:

bash peppercorn-v2/examples/serve.sh

Run two replicas on the same GPU:

REPLICAS=2 bash peppercorn-v2/examples/serve.sh

The server exposes an OpenAI-compatible API. The default model identifier is peppercorn.

Available environment variables include:

Variable Description
MODEL Hugging Face repository ID or local model directory
REPLICAS Number of serving replicas
PORT Base API port
MEM GPU memory utilization per replica
API_SERVERS Number of API server processes

The example configuration uses settings intended for page-level OCR, large request batches, and efficient asynchronous scheduling.

OpenAI-Compatible Client

The following example sends a document image to a running server.

import base64
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="none",
)

with open("page.png", "rb") as f:
    encoded_image = base64.b64encode(
        f.read()
    ).decode()

image_url = (
    "data:image/png;base64," + encoded_image
)

response = client.chat.completions.create(
    model="peppercorn",
    temperature=0,
    max_tokens=4096,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image_url",
                    "image_url": {
                        "url": image_url
                    },
                }
            ],
        }
    ],
)

print(response.choices[0].message.content)

Repeated Output and Loop Guard

Under greedy decoding, a small proportion of pages may generate repetitive text instead of terminating normally.

This behavior is more common with dense English documents, difficult born-digital pages, and mathematical content.

The provided client includes a loop guard that monitors streamed output for repeated spans or lines.

When repetition is detected, the client:

  1. Stops the generation request.
  2. Retains the output up to the first occurrence of the repeated content.
  3. Marks the result with a loop finish status.

In the v2 evaluation, the guard stopped 51 of 56 looping pages and incorrectly stopped 1 of the 1,936 pages that would otherwise have terminated normally.

These figures describe the tested evaluation set; actual rates depend on the workload.

Limitations

  • English OCR, mathematical expressions, and some real financial tables remain weaker than v1 on the reported evaluation sets.
  • Without the loop guard, approximately 2.8% of pages in the evaluation suite loop to the token limit.
  • The model is evaluated with greedy decoding. Sampling can increase the probability of repetitive output.
  • Images are processed with a maximum longest edge of 2,048 pixels. Small text on large-format documents may lose detail after resizing.
  • OCR accuracy depends on image quality, typography, layout complexity, and document content.
  • Structured output should be validated when exact transcription or preservation of table structure is required.

Repository Contents

The repository provides the model files and the components required for supported inference workflows.

Component Purpose
Model checkpoint and configuration Model weights and runtime configuration
Tokenizer and processor Text tokenization and image preprocessing
vllm_plugin/ Integration for vLLM inference
examples/serve.sh API server startup script
examples/client.py Streaming inference client with loop detection
examples/loop_guard.py Repeated-output detection
examples/offline_infer.py Offline inference

Training datasets, training scripts, and internal training implementation details are not documented as part of the public inference interface.

License and Usage Restrictions

Research and evaluation use only.

Peppercorn and its associated repository materials are provided subject to the applicable license terms.

Unless explicitly authorized in writing by the relevant rights holder, no permission is granted to copy, reproduce, modify, redistribute, sublicense, publish, or commercially exploit the model weights, source code, integration code, or other protected materials.

Access to a public repository does not, by itself, grant permission for uses prohibited by the applicable license terms.

Third-party components and materials may be subject to separate license conditions. Those conditions remain applicable and are not superseded by this notice.

Users must review the relevant license and attribution requirements before using or redistributing the model or associated materials.


Peppercorn — compact document OCR for Vietnamese and English.

Downloads last month
9
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support