Instructions to use DuyTa/Peppercorn-0.3b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DuyTa/Peppercorn-0.3b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="DuyTa/Peppercorn-0.3b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSeq2SeqLM processor = AutoProcessor.from_pretrained("DuyTa/Peppercorn-0.3b") model = AutoModelForSeq2SeqLM.from_pretrained("DuyTa/Peppercorn-0.3b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use DuyTa/Peppercorn-0.3b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "DuyTa/Peppercorn-0.3b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DuyTa/Peppercorn-0.3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/DuyTa/Peppercorn-0.3b
- SGLang
How to use DuyTa/Peppercorn-0.3b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "DuyTa/Peppercorn-0.3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DuyTa/Peppercorn-0.3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "DuyTa/Peppercorn-0.3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DuyTa/Peppercorn-0.3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use DuyTa/Peppercorn-0.3b with Docker Model Runner:
docker model run hf.co/DuyTa/Peppercorn-0.3b
Peppercorn (Hạt Tiêu), 0.3B
Peppercorn is a compact vision-language model for Vietnamese and English document OCR, with 328M parameters.
It takes a document page as input and produces structured Markdown, including HTML tables and LaTeX mathematical expressions.
Its name comes from the Vietnamese saying *nhỏ nhưng có võ — small, but mighty: compact in size, yet capable.
Compared with original design, Peppercorn uses approximately 4× fewer image tokens per page (~900 vs. ~3,600) and achieves higher throughput under our tested serving configuration.
It performs particularly well on Vietnamese administrative documents and synthetic tables, while English OCR, mathematical expressions, and some real-world financial tables remain areas for improvement.
Highlights
- 328M parameters for efficient document understanding.
- Vietnamese and English OCR, covering both born-digital documents and scanned pages.
- Structured output in Markdown, HTML tables, and LaTeX.
- Efficient visual encoding, using approximately 880 image tokens per page on average.
- High-throughput inference with vLLM.
- OpenAI-compatible API integration through the provided serving interface.
Intended Use
Peppercorn is designed for document digitization and structured text extraction, including:
- Vietnamese administrative documents.
- Scanned documents and born-digital pages.
- Documents containing tables and structured layouts.
- Mathematical expressions and technical documents.
- Batch document processing pipelines.
The model generates text from document images. Its output should be validated before being used in workflows where transcription errors could have significant consequences.
Architecture
Peppercorn uses a compact vision-language architecture designed specifically for efficient document OCR.
| Component | Description |
|---|---|
| Vision encoder | Extracts visual features from document images |
| Visual processing | Image tiling and feature aggregation for page-level understanding |
| Connector | Converts visual features into representations suitable for the language model |
| Text decoder | Generates structured text from visual input |
| Output formats | Markdown, HTML tables, and LaTeX |
Image Processing
The longest image edge is resized to a maximum of 2,048 pixels. Larger pages are processed using image tiles, together with a global view when applicable.
This allows the model to process full document pages while controlling the number of visual tokens.
The average input is approximately 880 image tokens per page, although the actual number depends on the document dimensions and tiling configuration.
Output Format
The model is intended to produce readable, structured text rather than unformatted character sequences.
For example, tables may be represented as HTML:
<table>
<tr>
<th>Item</th>
<th>Value</th>
</tr>
<tr>
<td>Example</td>
<td>123</td>
</tr>
</table>
Mathematical expressions may be represented using LaTeX:
\frac{a+b}{c}
The exact output structure depends on the document content and model generation behavior.
Evaluation
All CER values below are calculated on normalized text. Lower is better.
Evaluation uses greedy decoding with a maximum of 4,096 generated tokens per page unless otherwise specified.
| Evaluation Set | Pages | v1 Mean CER | v2 Mean CER | v2 Median CER | v2 Looped Pages |
|---|---|---|---|---|---|
| Vietnamese born-digital | 447 | 0.038 | 0.046 | 0.004 | 9 |
| Vietnamese born-digital, hard | 129 | 0.139 | 0.140 | 0.032 | 13 |
| Vietnamese scans | 189 | 0.062 | 0.071 | 0.011 | 7 |
| English pages | 280 | 0.127 | 0.140 | 0.032 | 14 |
| English crops | 100 | 0.013 | 0.027 | 0.000 | 1 |
| Vietnamese crops | 200 | 0.018 | 0.022 | 0.000 | 2 |
| Vietnamese mathematics | 162 | 0.077 | 0.088 | 0.023 | 7 |
| Synthetic Vietnamese tables | 400 | 0.023 | 0.014 | 0.000 | 1 |
| Real financial tables (HTML) | 85 | 0.082 | 0.106 | 0.047 | 2 |
| All evaluation sets | 1,992 | 0.058 | 0.064 | 0.006 | 56 (2.8%) |
| Administrative documents | 732 | 0.0224 | 0.0169 | — | — |
| Administrative pages containing tables | 141 | 0.057 | 0.043 | — | — |
Interpretation
Peppercorn improves on the administrative document evaluation set and synthetic table recognition.
Across the complete evaluation suite, v1 retains a lower mean CER. The results indicate that the newer model's efficiency improvements do not translate into uniform accuracy gains across every document category.
A page is classified as looped if generation does not terminate within the token limit or produces output exceeding three times the reference length.
With the loop guard enabled, the v2 evaluation CER decreases to 0.061.
On pages where neither model loops, v2 performs comparably to v1.
These results are specific to the evaluation datasets and inference settings described here. They should not be interpreted as a guarantee of accuracy on arbitrary documents.
Throughput
Throughput was measured on an RTX PRO 6000 (96 GB) using vLLM 0.29 and a workload of 800 mixed Vietnamese document pages, including born-digital pages, scans, and tables.
The serving setup used the provided example scripts, streaming inference, and 1,024 concurrent requests.
| Configuration | Throughput | Output Tokens per Page |
|---|---|---|
| 1 engine, 8 API servers, no loop guard | 14.9 pages/s | 687 |
| 1 engine, 8 API servers, loop guard enabled | 15.1 pages/s | 610 |
| 2 engines on one GPU, 4 API servers each, loop guard enabled | 22.1 pages/s | 602 |
A single engine does not fully utilize the GPU under this workload. Running two engine replicas on the same GPU increases aggregate throughput by approximately 1.4× in the tested configuration.
Streaming individual tokens introduces additional overhead. The recommended serving configuration uses 32-token streaming intervals to balance responsiveness and throughput.
Actual performance varies with image dimensions, document complexity, output length, concurrency, GPU configuration, and serving parameters.
Installation
Requirements
- Python
- PyTorch
transformers>=5for Transformers inferencevllm>=0.29for the provided vLLM integration
Download the Model
pip install "vllm>=0.29"
hf download DuyTa/peppercorn-v2-0.3b \
--local-dir peppercorn-v2
Install the provided vLLM integration:
pip install ./peppercorn-v2/vllm_plugin
The integration registers the model architecture with vLLM through its plugin mechanism.
Inference with Transformers
The following example loads the model and extracts text from a document image.
from transformers import AutoModelForImageTextToText, AutoProcessor
from PIL import Image
import torch
repo = "DuyTa/peppercorn-v2-0.3b"
processor = AutoProcessor.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(
repo,
dtype=torch.bfloat16,
).cuda().eval()
image = Image.open("page.png").convert("RGB")
prompt = processor.apply_chat_template(
[
{
"role": "user",
"content": [{"type": "image"}],
}
],
add_generation_prompt=True,
)
inputs = processor(
text=prompt,
images=[[image]],
return_tensors="pt",
).to("cuda")
inputs["pixel_values"] = inputs["pixel_values"].to(
torch.bfloat16
)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=4096,
do_sample=False,
)
generated_tokens = output[
0,
inputs["input_ids"].shape[1]:,
]
print(
processor.decode(
generated_tokens,
skip_special_tokens=True,
)
)
The prompt contains the document image without additional textual instructions.
A pad_token_id warning referencing 128002 may appear because of the default configuration inherited from the Idefics3 configuration. In the described setup, this warning is harmless.
Inference with vLLM
The provided vLLM integration supports offline inference and OpenAI-compatible serving.
Offline Inference
python peppercorn-v2/examples/offline_infer.py \
--model DuyTa/peppercorn-v2-0.3b \
page1.png page2.jpg
The script prints the extracted Markdown along with the generation finish reason.
Start the API Server
Run one replica:
bash peppercorn-v2/examples/serve.sh
Run two replicas on the same GPU:
REPLICAS=2 bash peppercorn-v2/examples/serve.sh
The server exposes an OpenAI-compatible API. The default model identifier is peppercorn.
Available environment variables include:
| Variable | Description |
|---|---|
MODEL |
Hugging Face repository ID or local model directory |
REPLICAS |
Number of serving replicas |
PORT |
Base API port |
MEM |
GPU memory utilization per replica |
API_SERVERS |
Number of API server processes |
The example configuration uses settings intended for page-level OCR, large request batches, and efficient asynchronous scheduling.
OpenAI-Compatible Client
The following example sends a document image to a running server.
import base64
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="none",
)
with open("page.png", "rb") as f:
encoded_image = base64.b64encode(
f.read()
).decode()
image_url = (
"data:image/png;base64," + encoded_image
)
response = client.chat.completions.create(
model="peppercorn",
temperature=0,
max_tokens=4096,
messages=[
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": image_url
},
}
],
}
],
)
print(response.choices[0].message.content)
Repeated Output and Loop Guard
Under greedy decoding, a small proportion of pages may generate repetitive text instead of terminating normally.
This behavior is more common with dense English documents, difficult born-digital pages, and mathematical content.
The provided client includes a loop guard that monitors streamed output for repeated spans or lines.
When repetition is detected, the client:
- Stops the generation request.
- Retains the output up to the first occurrence of the repeated content.
- Marks the result with a
loopfinish status.
In the v2 evaluation, the guard stopped 51 of 56 looping pages and incorrectly stopped 1 of the 1,936 pages that would otherwise have terminated normally.
These figures describe the tested evaluation set; actual rates depend on the workload.
Limitations
- English OCR, mathematical expressions, and some real financial tables remain weaker than v1 on the reported evaluation sets.
- Without the loop guard, approximately 2.8% of pages in the evaluation suite loop to the token limit.
- The model is evaluated with greedy decoding. Sampling can increase the probability of repetitive output.
- Images are processed with a maximum longest edge of 2,048 pixels. Small text on large-format documents may lose detail after resizing.
- OCR accuracy depends on image quality, typography, layout complexity, and document content.
- Structured output should be validated when exact transcription or preservation of table structure is required.
Repository Contents
The repository provides the model files and the components required for supported inference workflows.
| Component | Purpose |
|---|---|
| Model checkpoint and configuration | Model weights and runtime configuration |
| Tokenizer and processor | Text tokenization and image preprocessing |
vllm_plugin/ |
Integration for vLLM inference |
examples/serve.sh |
API server startup script |
examples/client.py |
Streaming inference client with loop detection |
examples/loop_guard.py |
Repeated-output detection |
examples/offline_infer.py |
Offline inference |
Training datasets, training scripts, and internal training implementation details are not documented as part of the public inference interface.
License and Usage Restrictions
Research and evaluation use only.
Peppercorn and its associated repository materials are provided subject to the applicable license terms.
Unless explicitly authorized in writing by the relevant rights holder, no permission is granted to copy, reproduce, modify, redistribute, sublicense, publish, or commercially exploit the model weights, source code, integration code, or other protected materials.
Access to a public repository does not, by itself, grant permission for uses prohibited by the applicable license terms.
Third-party components and materials may be subject to separate license conditions. Those conditions remain applicable and are not superseded by this notice.
Users must review the relevant license and attribution requirements before using or redistributing the model or associated materials.
Peppercorn — compact document OCR for Vietnamese and English.
- Downloads last month
- 9