source: Hugging Face Blog: NeoMME: an efficient Multimodal-native and Multilingual Encoder
level: technical
neomme is a new family of multilingual multimodal encoders in two sizes, 260m and 800m parameters. unlike many visual language models, it does not use a separate pretrained vision tower or causal language model. a single bidirectional transformer processes both text tokens and raw image patches. the model is trained from scratch using a masked discrete-diffusion objective. it is fine-tuned for visual document retrieval using colpali's page-image approach, returning dense and late-interaction embeddings in one forward pass.
on the vidore v3 benchmark, neomme-retriever-260m reaches 0.523 ndcg@10, within 0.002 of colqwen2.5 while using about 14 times fewer parameters. the 800m model reaches 0.556, close to vultron retriever flash. both models lie on the model-size pareto frontier. at 2048x2048 input on an nvidia l40s gpu, the 260m model encodes about 51 pages per second, nearly twice colmodernvbert's throughput. hierarchical token pooling and asymmetric quantization reduce late-interaction index storage from about 1.5 mb to 6 kb per page, a 255-fold reduction, while retaining over 95% of baseline ndcg@10.
neomme's architecture uses dynamic image resolution, long bidirectional context of 16,384 tokens, and modern encoder improvements like grouped-query attention and 2d rotary position embeddings. pretraining mixes multilingual text, code, mathematics, and images, processing about 524 billion tokens. the model is available in hugging face transformers under apache 2.0 license. this design removes the overhead of generative visual language models, making it practical for high-resolution document retrieval at scale.
why it matters: neomme offers a compact and fast way to build visual document retrieval systems, reducing compute and storage costs while maintaining competitive accuracy.
source: Hugging Face Blog: NeoMME: an efficient Multimodal-native and Multilingual Encoder