source: Hugging Face Blog: Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
level: technical
Sentence Transformers v6.0 adds MultiVectorEncoder, a new model type for ColBERT-style late interaction retrieval. It loads PyLate and Stanford-NLP ColBERT checkpoints directly, and ColPali models for visual document retrieval through the same API. Unlike dense models that compress text into one vector, multi-vector models keep one vector per token and score with MaxSim, preserving token-level matches. This improves retrieval for exact identifiers and multi-requirement queries, at the cost of larger indexes.
Encoding 4,874 Natural Questions passages with lightonai/LateOn produced 608,414 token vectors, averaging 124.8 per passage. That is 311.5 MB in float32, about 42 times the storage of a MiniLM dense index. However, compressed indexes like fast-plaid reduce this to 92 MB, comparable to dense models. The MaxSim operator sums each query token's highest similarity to any document token, enabling soft alignment. For example, 'live' matches 'inhabit' at 0.94 despite no shared characters.
Multi-vector models are asymmetric: queries and documents use different prefixes, length caps, and masks. Users must call encode_query and encode_document separately. Checkpoints configure their own settings, such as ColBERTv2 padding queries to 32 tokens and truncating documents at 180. Document length caps can be lifted per call, but this increases index size. The library requires transformers v5.x, torch 2.2+, and huggingface-hub v1.x, so users should plan upgrades before installing.
why it matters: Multi-vector retrieval improves accuracy for exact matches and complex queries, but requires managing larger indexes and new encoding workflows.
source: Hugging Face Blog: Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers