⚡
Brief IA
›

Multi-vectors: Unsupervised Checkpoints Adapt Better

🔬 Research·Tom Levy·

Multi-vectors: Unsupervised Checkpoints Adapt Better

Multi-vectors: Unsupervised Checkpoints Adapt Better
⚡
Key Takeaways
1Unsupervised checkpoints adapted better to the domain than fully trained starting points in comparisons on MIRIAD.
2Truncation on long passages cost up to 0.24 NDCG@10 for texts averaging 941 tokens.
3MultiVectorEncoder allows for fine-tuning an existing checkpoint or adding a fresh head to a backbone with datasets and the Hub.
💡Why it matters — The choice of starting point and length management strongly influence the measured retrieval quality, without changing the architecture.

Experiments on medical games highlight an adaptation gain in the domain for checkpoints labeled -unsupervised. Explicit management of document lengths avoids NDCG@10 losses measured up to 0.24 when passages exceed 512 tokens. MultiVectorEncoder allows for fine-tuning an existing checkpoint or adding a fresh head to a backbone, utilizing tools based on datasets and the Hugging Face Hub.

Starting with an -unsupervised checkpoint, according to comparative trials

Tests conducted on six initial configurations, all trained under the same protocol on 25,000 medical question-passage pairs and then evaluated with 1,000 questions and 50,000 passages, show that checkpoints identified by the suffix -unsupervised exhibit a significantly better adaptation capacity to an unseen domain than those that are fully finalized. This observation has been confirmed for two distinct groups of models. Finalized checkpoints showed little progress, or even a decrease in performance, regardless of the learning rates examined. The proposed explanation is that their effectiveness stems from the fact that they intervene just after a massive contrastive pre-training, but before a supervised fine-tuning on general retrieval, which would help preserve this adjustment during subsequent training in the targeted domain. In practice, it is recommended to prioritize a pre-supervised checkpoint when available. If not, it is advised to use a new projection head associated with a robust pre-trained retrieval backbone, while a completely finalized starting point proves to be the least relevant for adaptation to a new domain. Furthermore, a newly initialized head on Alibaba-NLP/gte-modernbert-base achieved in the reported experiments a performance only 0.03 points below some existing checkpoints with 25,000 training pairs.

Long documents: check and lift costly truncations

The standard configuration of many retrieval models targets short passages: ColBERT checkpoints frequently truncate at 180 or 300 tokens, popular dense models at 256 or 512, consistent with MS MARCO-type data rarely exceeding these lengths. On long documents, these limits lead to silently ignoring a large portion of the text before scoring. In a medical evaluation where passages averaged 941 tokens, truncation cost up to 0.24 NDCG@10, more than the gap between architectures. For fine-tuning, it is advisable to first examine lengths: the cited medical passages go up to 1,400 tokens, and the mLateOn family can serve a complete context of 8,192 tokens. If the starting point limits length, it is possible to disable task ceilings and revert to the maximum length of the tokenizer, to be configured upon loading. An example of initialization explicitly sets a maximum length of 8,192 tokens to avoid unwanted truncation.

⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

Two training paths: fine-tune a checkpoint or add a fresh head

Adapting an existing multi-vector model does not require modifying its architecture: each checkpoint includes its own markers for queries and documents, a projection head, as well as a scoring jump list, which is generally wise to retain by only changing the elements required by the data. It is also possible to configure MultiVectorEncoder to directly use a base transformer: in this case, a token-level projection is added and initialized randomly, which implies prior training before any use. This method is also compatible with dense embedding backbones that offer good performance.

How late interaction works in this pipeline

Late interaction retains a vector per token and evaluates similarity via MaxSim, with each token of the query finding its best match in the document before summing the scores. The ColBERT pipeline relies on a Transformer that generates contextualized embeddings, a dense projection reducing each token to 128 dimensions, a MultiVectorMask selecting the tokens contributing to the scoring, and normalization at the token level. In practice, this fine matching retains signals that aggregation into a single vector tends to smooth out, and it performs well even with little domain fine-tuning data, at the cost of larger indexes.

Data and training tooling with datasets and the Hub

A MultiVectorEncoder training session combines a model, a dataset, a loss function, training arguments, and optional evaluators within a trainer. MultiVectorEncoderTrainer relies on datasets.Dataset or datasets.DatasetDict objects for training and evaluation. The data can come from the Hugging Face Datasets Hub or local sources in CSV, JSON, Parquet, Arrow, or SQL. Many datasets directly usable with Sentence Transformers are tagged "sentence-transformers" on the Hub for easy retrieval. When building custom training, the length of documents can be defined according to the specific needs of the corpus.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.