
TL;DR
We are introducing NeoMME, a multilingual, multimodal encoder family with 260 million- and 800 million-parameter models. Unlike many generative vision-language models, NeoMME does not use a separate pretrained vision tower or a causal language model. A single bidirectional Transformer processes text tokens and raw image patches together, and the entire model is trained from scratch with a masked discrete diffusion objective.
We fine-tuned NeoMME for visual document retrieval using ColPali’s page-image approach. NeoMME-Retriever returns both dense and late-interaction embeddings in a single forward pass. Both model sizes lie on the Pareto frontier of nDCG@10 versus model size on ViDoRe v3. When evaluated on an NVIDIA L40S GPU with matched 2048×2048 image inputs, the 260 million-parameter model encodes approximately 51 pages per second—about twice the throughput of ColModernVBERT. Hierarchical token pooling and asymmetric quantization reduce late-interaction index storage from approximately 1.5 MB per page to 6 kB—a 255× reduction—while retaining more than 95% of the baseline nDCG@10.
NeoMME is available in Hugging Face Transformers. We are releasing all model checkpoints under the Apache 2.0 license.
Why Do We Need Yet Another Multimodal Encoder?
Many recent visual document retrievers are adapted from pretrained generative vision-language models. A separately pretrained vision encoder generates visual features, which a projector then maps into the language model’s input space. A causal decoder subsequently processes the combined image and text representations. Retrieval, classification, and token labeling do not require autoregressive text generation, so they do not need a causal decoder or the parameter and computational overhead introduced by this architecture.
ModernBERT brought efficient architectural and training improvements to bidirectional encoders. In visual document retrieval, ModernVBERT uses a bidirectional ModernBERT-style text encoder while retaining a separately pretrained SigLIP2 vision tower. We wanted to build on this work by designing and training a multimodal encoder that does not inherit the parameter and computational overhead of vision-language models.
NeoMME (pronounced “nee-oh-me,” IPA /ˈniː.oʊ.mi/) is a multilingual, multimodal foundation encoder that uses a single Transformer encoder to generate vector representations for input text and/or images. It is not based on an existing pretrained vision tower, text encoder, or text decoder.

Unlike dual-tower encoders and vision-language models, NeoMME processes image patches and text tokens within a single bidirectional Transformer, without using a pretrained vision tower, pretrained text encoder, or pretrained text decoder.
Because images and text use the same computational path, NeoMME can more easily support pretraining, fine-tuning, parallelization, and serving across both modalities.
The NeoMME Encoder Backbone
One Transformer for Images and Text
NeoMME is available in two model sizes: 260 million parameters and 800 million parameters. Both versions use the same architecture:
-
Native multimodal inputs: Text inputs use factorized token embeddings, while images are divided into non-overlapping 32×32 patches and projected through a small MLP. Both then enter the same Transformer encoder.
-
Dynamic image resolution: Images retain their aspect ratios and dimensions. This allows the model to use more tokens for high-resolution, information-dense document pages and fewer tokens for smaller images with less content.
-
Long-range bidirectional context: Both models have a context length of 16,384 tokens, enough to accommodate up to two standard 3840×2160 4K UHD images. Most layers use symmetric sliding-window attention, while one layer in every six and the final layer use global attention.
-
Modern encoder stack: NeoMME incorporates recent encoder improvements, including grouped-query attention, query-key normalization, gated attention, 2D rotary position embeddings, and squared ReLU MLPs.
-
Multilingual text: We trained a 131k-vocabulary BPE tokenizer from scratch on multilingual text, code, mathematics, and machine-generated image transcriptions.

Alternating sliding-window attention and global attention layers in the NeoMME encoder stack.
Learning from Images Through Masked Text
We pretrained NeoMME from scratch as a discrete masked diffusion text denoiser. For each plain-text sample, we uniformly sample a corruption rate between 0 and 1. Each eligible text token is then independently masked at that rate.
Multimodal samples use corruption rates between 0.3 and 1. Image patches remain visible while NeoMME reconstructs the masked text. At lower masking rates, the model can often recover missing words from the surrounding text alone. For example, even without an image, filling in “cat” in The [MASK] sat on the mat is reasonable. However, higher masking rates force the model to learn image-based descriptions when the unmasked input text tokens provide little information.

Higher text corruption rates eliminate language-only shortcuts and encourage NeoMME to use visible image evidence.
The pretraining data combines multilingual text, code, mathematics, natural images, and document images. Each model processes approximately 524 billion packed input tokens, including 290 billion tokens from plain-text samples. Compared with ModernBERT’s 2 trillion-token training budget, this is a relatively small amount of text data. We therefore chose the NorMuon optimizer to improve data efficiency during training.
NeoMME-Retriever
To meaningfully evaluate the model backbone downstream, we fine-tuned NeoMME for visual document retrieval using the page-image approach introduced by ColPali. Traditional text-based retrieval typically retrieves text chunks, whereas NeoMME-Retriever ranks screenshots of document pages, bypassing the full OCR preprocessing pipeline required to extract text from PDFs. Treating pages as images preserves layout, charts, tables, font styles and sizes, and other visual cues that even a perfect OCR model cannot capture.
A Dual-Head Design for Dense and Late-Interaction Retrieval
NeoMME-Retriever reuses the NeoMME backbone.