Tom Aarsen's picture
🏗️ Building on HF

Tom Aarsen

tomaarsen
huggingface

AI & ML interests

NLP: text embeddings, information retrieval, named entity recognition, few-shot text classification

Recent Activity

posted an update 28 minutes ago
🚨 I've just published Sentence Transformers v6.0, introducing MultiVectorEncoder: ColBERT-style late interaction models are now a fourth model type, for training, inference, and interpretation, alongside the dense, sparse, and reranker models! Details: Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator. That preserves token-level matching information that a single vector has to average away. It is also the state of the art for visual document retrieval, where a text query is matched against page images directly, charts and tables included, with no OCR step in between. Any PyLate, Stanford ColBERT, or ColPali checkpoint loads straight into the same familiar API: model.encode_query(), model.encode_document(), and model.similarity() just work, whether the documents are texts or page images. Does it help? LightOn trained LateOn (multi-vector) and DenseOn (dense) on the same data with the same 149M ModernBERT backbone, and the multi-vector model wins on 9 of the 13 NanoBEIR datasets: 0.6868 vs 0.6764 mean NDCG@10. The price is a bigger index, and the new HierarchicalTokenPooling module halves it at roughly no retrieval cost. Antoine Chaffin, Raphaël Sourty, and I wrote a blog post walking through multi-vector models in practice: loading the various checkpoint formats, encoding and scoring, plugging them into a search stack, running them on page images, and keeping the index affordable. Check it out if you want to get started, or just point your Agent to the URL: https://huggingface.co/blog/multi-vector-encoder pip install sentence-transformers==6.0.0 Release notes: https://github.com/huggingface/sentence-transformers/releases/tag/v6.0.0
View all activity

Organizations

Hugging Face's profile picture Sentence Transformers's profile picture Sentence Transformers - Cross-Encoders's profile picture Hugging Face Internal Testing Organization's profile picture SetFit's profile picture Massive Text Embedding Benchmark's profile picture Hugging Face Fellows's profile picture Nomic AI's profile picture Open-Source AI Meetup's profile picture Hugging Face OSS Metrics's profile picture Blog-explorers's profile picture Sentence Transformers Testing's profile picture mLLM multilingual's profile picture Social Post Explorers's profile picture Answer.AI's profile picture gg-tt's profile picture Distillation Hugs's profile picture Hugging Face Discord Community's profile picture Perplexity's profile picture Bert ... but new's profile picture EuroBERT's profile picture Sentence Transformers - Cross-Encoders Testing's profile picture hf-inference's profile picture Transformers Community's profile picture gg-hf-gm's profile picture Hugging Face Context Course's profile picture Sentence Transformers - Sparse Encoders's profile picture Sentence Transformers - Sparse Encoders Testing's profile picture Late Interaction's profile picture hfpp's profile picture MongoDB AI Community's profile picture ML intern explorers's profile picture Humanity's Last Hackathon's profile picture Hugging Apps's profile picture Sentence Transformers - Multi Vector Encoders's profile picture Sentence Transformers - Multi Vector Encoders Testing's profile picture