GELATO qwen document retriever

Complete inference checkpoint: frozen Qwen3.5 vision tower, frozen Qwen/Qwen3-Embedding-0.6B text backbone, and a learned projector plus image delimiters. Outputs normalized 1024-dimensional vectors for text queries and page images. Includes 699,518,208 model parameters. No training code, datasets or caches.

Checkpoint 12,300, selected by the best completed public ViDoRe v3 score: 35.72 macro nDCG@10 across eight corpora and six query languages. This benchmark was used for selection; it is not an untouched test. Trained on approximately 273k multilingual VDR positive pairs, with frozen backbones and Matryoshka contrastive training. No retraining was performed. Frozen backbone weights use the BF16 values used during training and evaluation. Backbone pins are recorded in config.json.

Requires Python 3.12, Transformers >=5.16.1, Sentence Transformers >=6.0.1, PyTorch, Pillow and safetensors. BF16 CUDA inference matches the evaluated recipe. Read the included small inference files before allowing custom code.

Sentence Transformers

from PIL import Image
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("AdrienB134/gelato-qwen", trust_remote_code=True, device="cuda")
queries = model.encode_query(["Find the page showing annual revenue."], convert_to_tensor=True)
pages = model.encode([Image.open("page.png")], convert_to_tensor=True, batch_size=1)
scores = queries @ pages.T

encode() accepts homogeneous batches of strings or PIL images. Text defaults to the retrieval query prompt. encode_query() uses the query prompt; encode_document() uses the frozen text backbone's document prompt. Images use the trained multimodal path. Mixed text/image batches and video are not supported. Start with a small image batch and increase it according to GPU memory.

Hugging Face Transformers

from PIL import Image
import torch
from transformers import AutoModel, AutoProcessor

repo = "AdrienB134/gelato-qwen"
processor = AutoProcessor.from_pretrained(repo, trust_remote_code=True)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, dtype=torch.bfloat16).to("cuda").eval()
text = processor(text=["Find the page showing annual revenue."]).to("cuda")
image = processor(images=[Image.open("page.png")]).to("cuda")
with torch.inference_mode():
    queries = model(**text).embeddings
    pages = model(**image).embeddings
scores = queries @ pages.T

All weights, tokenizer and image processor are bundled. After downloading, both interfaces load from the local directory with local_files_only=True and Hub offline mode. No GELATO package or original backbone repository is required.

Images preserve aspect ratio within the original 262,144–1,310,720 pixel budget (256–1,280 merged visual tokens). Text pooling, image-sequence framing, prompts and normalization preserve the evaluated checkpoint. The unused language model inside the original Qwen3.5 vision checkpoint is omitted; all vision-tower weights are preserved. Source authors and licenses are documented in NOTICE.

Downloads last month
18
Safetensors
Model size
0.7B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AdrienB134/gelato-qwen

Finetuned
(295)
this model