Instructions to use AdrienB134/gelato-qwen with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use AdrienB134/gelato-qwen with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("AdrienB134/gelato-qwen", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
GELATO qwen document retriever
Complete inference checkpoint: frozen Qwen3.5 vision tower, frozen
Qwen/Qwen3-Embedding-0.6B text backbone, and a learned projector plus image delimiters.
Outputs normalized 1024-dimensional vectors for text queries and page images.
Includes 699,518,208 model parameters. No training code, datasets or caches.
Checkpoint 12,300, selected by the best completed public ViDoRe v3
score: 35.72 macro nDCG@10 across eight corpora and six query
languages. This benchmark was used for selection; it is not an untouched test.
Trained on approximately 273k multilingual VDR positive pairs, with frozen
backbones and Matryoshka contrastive training. No retraining was performed.
Frozen backbone weights use the BF16 values used during training and evaluation.
Backbone pins are recorded in config.json.
Requires Python 3.12, Transformers >=5.16.1, Sentence Transformers >=6.0.1, PyTorch, Pillow and safetensors. BF16 CUDA inference matches the evaluated recipe. Read the included small inference files before allowing custom code.
Sentence Transformers
from PIL import Image
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("AdrienB134/gelato-qwen", trust_remote_code=True, device="cuda")
queries = model.encode_query(["Find the page showing annual revenue."], convert_to_tensor=True)
pages = model.encode([Image.open("page.png")], convert_to_tensor=True, batch_size=1)
scores = queries @ pages.T
encode() accepts homogeneous batches of strings or PIL images. Text defaults to
the retrieval query prompt. encode_query() uses the query prompt;
encode_document() uses the frozen text backbone's document prompt. Images use
the trained multimodal path. Mixed text/image batches and video are not supported.
Start with a small image batch and increase it according to GPU memory.
Hugging Face Transformers
from PIL import Image
import torch
from transformers import AutoModel, AutoProcessor
repo = "AdrienB134/gelato-qwen"
processor = AutoProcessor.from_pretrained(repo, trust_remote_code=True)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, dtype=torch.bfloat16).to("cuda").eval()
text = processor(text=["Find the page showing annual revenue."]).to("cuda")
image = processor(images=[Image.open("page.png")]).to("cuda")
with torch.inference_mode():
queries = model(**text).embeddings
pages = model(**image).embeddings
scores = queries @ pages.T
All weights, tokenizer and image processor are bundled. After downloading, both
interfaces load from the local directory with local_files_only=True and Hub
offline mode. No GELATO package or original backbone repository is required.
Images preserve aspect ratio within the original 262,144–1,310,720 pixel budget
(256–1,280 merged visual tokens). Text pooling, image-sequence framing, prompts
and normalization preserve the evaluated checkpoint. The unused language model
inside the original Qwen3.5 vision checkpoint is omitted; all vision-tower weights
are preserved. Source authors and licenses are documented in NOTICE.
- Downloads last month
- 18