Skip to content
Main Site News Console

Train and Fine-Tune Multi-Vector Embedding Models with Sentence Transformers

· Hugging Face Translated
教程模型卡

Sentence Transformers is a Python library for using and training embedding and reranking models across a wide range of applications, including retrieval-augmented generation, semantic search, and semantic text similarity. Its v6.0 release introduced a fourth model type, MultiVectorEncoder, for ColBERT-style late-interaction retrieval, along with a complete training workflow. In this article, I’ll show you how to fine-tune a multi-vector model so that it outperforms general-purpose retrievers on your data. The same approach can also be used to train powerful new multi-vector models from scratch. Everything described below can be run with pip install -U "sentence-transformers[train]".

Fine-tuning a multi-vector model involves several components: the model itself, the dataset, the loss function, training arguments, an evaluator, and a trainer class. I’ll introduce each of these components in turn and use practical examples to demonstrate how to fine-tune a high-performing multi-vector model.

Finally, in the evaluation section, I’ll show that the multi-vector-encoder/mLateOn-medical model, which I trained for 14.5 hours on a single RTX 3090 using the methods described in this article, easily outperformed every general-purpose retrieval model I could find in my medical retrieval benchmark—including dense, sparse, lexical, and multi-vector models.

NDCG@10 and effective parameter count on MIRIAD: fine-tuned mLateOn-medical achieves top-tier performance at a smaller scale, surpassing the strongest general-purpose models

If you’re interested in fine-tuning dense embedding models, sparse embedding models, or reranking models, see my previous articles on training and fine-tuning embedding models, training and fine-tuning sparse embedding models, and training and fine-tuning reranking models.

This article explains how to train multi-vector models. If you want to learn how to use multi-vector models—including loading, encoding, and indexing them in a vector database—see the companion article Multi-Vector (Late-Interaction) Embedding Models with Sentence Transformers.

What Are Multi-Vector Models?

Dense embedding models compress an entire piece of text into a single vector, and the similarity between two such summaries can be computed with a single dot product. Multi-vector models (also known as late-interaction models or ColBERT-style models) skip this compression step. Instead, they retain a small vector for each token and use the MaxSim operator to calculate the score between a query and a document: each query token finds the document token that matches it most closely, and those scores are then summed. Token-level matching preserves fine-grained signals that would otherwise have to be averaged into a single vector. This often leads to stronger retrieval performance, at the cost of a larger index.

The companion article Multi-Vector Embedding Models covers the architecture, encoding, scoring, and indexing in detail, so I won’t repeat that material here. Let’s move straight on to training.

Dense Embeddings and Multi-Vector Late Interaction

Why Fine-Tune?

Fine-tuning a multi-vector model on a specific domain can significantly improve its retrieval performance: web search, legal discovery, code search, and scientific literature review all differ in vocabulary, query style, and definitions of relevance. Because queries and documents are matched token by token, multi-vector models can capture fine-grained domain signals that single-vector models typically average out. Even a moderate amount of in-domain fine-tuning data can produce strong results.

In addition, most published retrieval models are configured for short text passages. Classic ColBERT checkpoints truncate documents to 180 or 300 tokens, while many popular dense models truncate them to 256 or 512 tokens because their MS MARCO-style training data rarely exceeds those lengths. If your documents are longer, these models silently discard much of their content before scoring. In my medical benchmark, passages averaged 941 tokens; I found that truncation alone could reduce NDCG@10 by as much as 0.24—far more than the differences between model architectures. When training your own model, you can configure the document length to suit your data.

LightOn encountered the same issue in code retrieval: the general-purpose LateOn model was not ideal, so they trained LateOn-Code. Whether your domain is medicine, law, finance, or internal company documentation, there is unlikely to be an official model built specifically for you. This article will show you how to build a model for your own domain in just a few hours using a consumer-grade GPU.

Training Components

Training a MultiVectorEncoder model involves the following components:

  • Model: The model to fine-tune, or the architecture to build from scratch.

  • Dataset: The data used for training and evaluation.

  • Loss function: The function used to measure model performance and guide the optimization process.

  • Training arguments (optional): Arguments that affect training performance, tracking, and debugging.

  • Evaluator (optional): The class used to evaluate the model before, during, or after training.

  • Trainer: The component that brings all the training components together.

Let’s take a closer look at each component.

Model

Multi-vector training gives you several possible starting points, and the choice of starting point may matter more than you expect.

Fine-Tuning an Existing Multi-Vector Model

If you want to further fine-tune an existing multi-vector model, you don’t need to worry about the architecture at all:

from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder(
"lightonai/mLateOn-unsupervised",
model_kwargs={"torch_dtype": "float32"},
processor_kwargs={"model_max_length": 8192}, # tokenizer 层面的词元数量限制
)

This checkpoint contains its