Back to Embedding Research

Modern Sparse Neural Retrieval: From DeepCT to SPLADE++

Evgeniya Sukhodolskaya

·

October 23, 2024

Modern Sparse Neural Retrieval: From DeepCT to SPLADE++

Finding enough time to study all the modern solutions while keeping your production running is rarely feasible. Dense retrievers, hybrid retrievers, late interaction… How do they work, and where do they fit best? If only we could compare retrievers as easily as products on Amazon!

We explored the most popular modern sparse neural retrieval models and broke them down for you. By the end of this article, you’ll have a clear understanding of the current landscape in sparse neural retrieval and how to navigate through complex, math-heavy research papers with sky-high NDCG scores without getting overwhelmed.

Sparse Neural Retrieval: As If Keyword-Based Retrievers Understood Meaning

Keyword-based (lexical) retrievers like BM25 provide a good explainability. If a document matches a query, it’s easy to understand why: query terms are present in the document, and if these are rare terms, they are more important for retrieval.

Lexical retrievalExact terms · word statistics
Query
WhatisQdrant

Match only terms present in the document.

Document 42
Qdrantisavectordatabase
Document term frequency (TF)
Qdrant1is1a1vector1database1
Word statistics → relevance score

Sum contributions from matching query terms. BM25 combines document TF with corpus statistics: inverse document frequency (IDF) and average document length.

IDF factor
log(N / df(t))
TF and length factor
(k₁ + 1) · tf(t, d)k₁ · (1 − b + b · dl(d)/dlavg) + tf(t, d)

BM25 = Σt ∈ q IDF factor × TF and length factor. N: corpus size; df(t): documents containing t; k₁ and b: parameters; dl(d): document length; dlavg: average document length.

Query and document meet through shared terms, not learned embeddings. BM25 is the illustrative scoring formula from the original figure.

With their mechanism of exact term matching, they are super fast at retrieval. A simple inverted index, which maps back from a term to a list of documents where this term occurs, saves time on checking millions of documents.

Inverted indexTerm → document IDs
Posting lists
Qdrant242
vector542
database24842

Look up a term to find the documents that contain it. Document 42 occurs in all three lists.

A term points directly to the documents containing it, so retrieval can skip documents without that term.

Lexical retrievers are still a strong baseline in retrieval tasks. However, by design, they’re unable to bridge vocabulary and semantic mismatch gaps. Imagine searching for a “tasty cheese” in an online store and not having a chance to get “Gouda” or “Brie” in your shopping basket.

Dense retrievers, based on machine learning models which encode documents and queries in dense vector representations, are capable of breaching this gap and finding you “a piece of Gouda”.

Dense retrievalSentence embeddings
Query
WhatisQdrant
Dense embedding model

Example: text-embedding-3-small

Sentence embedding
−0.220.190.17…−0.510.330.40.23−2.5
Document 42
Qdrantisavectordatabase
Dense embedding model

Example: text-embedding-3-small

Sentence embedding
0.2−0.91.7…9.5−0.010.3−0.42.32.5
Similarity → relevance score
q · d = Σi qi di

Compare the whole embeddings; exact word overlap is not required.

Both branches encode an entire text into a dense vector. The vector coordinates are illustrative.

However, explainability here suffers: why is this query representation close to this document representation? Why, searching for “cheese”, we’re also offered “mouse traps”? What does each number in this vector representation mean? Which one of them is capturing the cheesiness?

Without a solid understanding, balancing result quality and resource consumption becomes challenging. Since, hypothetically, any document could match a query, relying on an inverted index with exact matching isn’t feasible. This doesn’t mean dense retrievers are inherently slower. However, lexical retrieval has been around long enough to inspire several effective architectural choices, which are often worth reusing.

Sooner or later, there should have been somebody who would say, “Wait, but what if I want something timeproof like BM25 but with semantic understanding?”

Sparse Neural Retrieval Evolution

Imagine searching for a “flabbergasting murder” story. ”Flabbergasting” is a rarely used word, so a keyword-based retriever, for example, BM25, will assign huge importance to it. Consequently, there is a high chance that a text unrelated to any crimes but mentioning something “flabbergasting” will pop up in the top results.

What if we could instead of relying on term frequency in a document as a proxy of term’s importance as it happens in BM25, directly predict a term’s importance? The goal is for rare but non-impactful terms to be assigned a much smaller weight than important terms with the same frequency, while both would be equally treated in the BM25 scenario.

How can we determine if one term is more important than another? Word impact is related to its meaning, and its meaning can be derived from its context (words which surround this particular word). That’s how dense contextual embedding models come into the picture.

All the sparse retrievers are based on the idea of taking a model which produces contextual dense vector representations for terms and teaching it to produce sparse ones. Very often, Bidirectional Encoder Representations from the Transformers (BERT) is used as a base model, and a very simple trainable neural network is added on top of it to sparsify the representations out. Training this small neural network is usually done by sampling from the MS MARCO dataset a query, relevant and irrelevant to it documents and shifting the parameters of the neural network in the direction of relevancy.

The Pioneer Of Sparse Neural Retrieval

DeepCTContextual word weights
Query
WhatisQdrant
BERT tokenizer → context
WhatisQdrant

BERT context: ht ∈ ℝ768 per token

First subtoken only: Q → Qdrant

Linear regression → round to integer
y = x · w + b
Query word weights
What5is9Qdrant340
Document 42
Qdrantisavectordatabase
BERT tokenizer → context
Qdrantisavectordatabase

BERT context: ht ∈ ℝ768 per token

First subtoken only: Q → Qdrant

Linear regression → round to integer
y = x · w + b
Document word weights
Qdrant230is2a3vector109database105
Exact word matching → sum
Qdrant340 × 230
is9 × 2
Σ matched contributions = 78,218
Matching words contribute their learned weights. The score is illustrative; it is not a complete BM25 calculation.
The authors of one of the first sparse retrievers, the Deep Contextualized Term Weighting framework (DeepCT), predict an integer word’s impact value separately for each unique word in a document and a query. They use a linear regression model on top of the contextual representations produced by the basic BERT model, the model’s output is rounded.

When documents are uploaded into a database, the importance of words in a document is predicted by a trained linear regression model and stored in the inverted index in the same way as term frequencies in BM25 retrievers. Then, the retrieval process is identical to the BM25 one.

Why is DeepCT not a perfect solution? To train linear regression, the authors needed to provide the true value (ground truth) of each word’s importance so the model could “see” what the right answer should be. This score is hard to define in a way that it truly expresses the query-document relevancy. Which score should have the most relevant word to a query when this word is taken from a five-page document? The second relevant? The third?

Sparse Neural Retrieval on Relevance Objective

DeepImpactWord-level scalar impacts
Query
WhatisQdrant
Document 42
Qdrantisavectordatabase
BERT tokenizer → context
Qdrantisavectordatabase

BERT context: ht ∈ ℝ768 per token

First subtoken only: Q → Qdrant

Two-layer NN · 768D → 1 scalar
Document word impacts
Qdrant2.2is0.3a0.1vector1.9database1.5
Sum impacts of matching query words
Qdrant1 × 2.2
is1 × 0.3
Σ matched impacts = 2.5
The document is encoded; query words select document impacts. Only the first subtoken of each document word is used. Values are illustrative.
It’s much easier to define whether a document as a whole is relevant or irrelevant to a query. That’s why the DeepImpact Sparse Neural Retriever authors directly used the relevancy between a query and a document as a training objective. They take BERT’s contextualized embeddings of the document’s words, transform them through a simple 2-layer neural network in a single scalar score and sum these scores up for each word overlapping with a query. The training objective is to make this score reflect the relevance between the query and the document.

Why is DeepImpact not a perfect solution? When converting texts into dense vector representations, the BERT model does not work on a word level. Sometimes, it breaks the words into parts. For example, the word “vector” will be processed by BERT as one piece, but for some words that, for example, BERT hasn’t seen before, it is going to cut the word in pieces as “Qdrant” turns to “Q”, “#dra” and “#nt”

The DeepImpact model (like the DeepCT model) takes the first piece BERT produces for a word and discards the rest. However, what can one find searching for “Q” instead of “Qdrant”?

Know Thine Tokenization

TILDEv2Token-level scalar weights
Query
WhatisQdrant
BERT tokenizer
WhatisQdrant
Binary query · 30,522 vocabulary slots
What1is1Q1dra1nt1

Other slots: 0

Document 42
Qdrantisavectordatabase
BERT tokenizer → context
Qdrantisavectordatabase

BERT context: ht ∈ ℝ768 per token

Scalar projection · 768D → 1 value
Document token weights
Q1.8dra3nt0.9is0.1a0.2vector2database1.8
Exact token matching → sum
Q1 × 1.8
dra1 × 3
nt1 × 0.9
is1 × 0.1
Σ matched impacts = 5.8
Q, dra, and nt remain separate matching dimensions. The unexpanded query has 1s for its tokens and 0s elsewhere; all example weights are illustrative.
To solve the problems of DeepImpact’s architecture, the Term Independent Likelihood MoDEl (TILDEv2) model generates sparse encodings on a level of BERT’s representations, not on words level. Aside from that, its authors use the identical architecture to the DeepImpact model.

Why is TILDEv2 not a perfect solution? A single scalar importance score value might not be enough to capture all distinct meanings of a word. Homonyms (pizza, cocktail, flower, and female name “Margherita”) are one of the troublemakers in information retrieval.

Sparse Neural Retriever Which Understood Homonyms

COILContextual token vectors
Query
WhatisQdrant
BERT tokenizer → context
WhatisQdrant

BERT context: ht ∈ ℝ768 per token

Linear projection · 768D → 32D
zt = W · ht + b
Query token vectors · 32D each
Whatq_Whatisq_isQq_Qdraq_drantq_nt
Document 42
Qdrantisavectordatabase
BERT tokenizer → context
Qdrantisavectordatabase

BERT context: ht ∈ ℝ768 per token

Linear projection · 768D → 32D
zt = W · ht + b
Document token vectors · 32D each
Qd_Qdrad_drantd_ntisd_isad_avectord_vectordatabased_database
Token-gated late interaction
1 · Exact token gate
isQdrant
2 · Dot product → max
mt = maxj: dⱼ = t qt · dj
3 · Sum
mis + mQ + mdra + mnt
Match token identities, compare their 32D contextual vectors, keep the best repeated occurrence, then sum. For bank, financial and river occurrences can score differently.

If one value for the term importance score is insufficient, we could describe the term’s importance in a vector form! Authors of the COntextualized Inverted List (COIL) model based their work on this idea. Instead of squeezing 768-dimensional BERT’s contextualised embeddings into one value, they down-project them (through the similar “relevance” training objective) to 32 dimensions. Moreover, not to miss a detail, they also encode the query terms as vectors.

For each vector representing a query token, COIL finds the closest match (using the maximum dot product) vector of the same token in a document. So, for example, if we are searching for “Revolut bank <finance institution>” and a document in a database has the sentence “Vivid bank <finance institution> was moved to the bank of Amstel <river>”, out of two “banks”, the first one will have a bigger value of a dot product with a “bank” in the query, and it will count towards the final score. The final relevancy score of a document is a sum of scores of query terms matched.

Why is COIL not a perfect solution? This way of defining the importance score captures deeper semantics; more meaning comes with more values used to describe it. However, storing 32-dimensional vectors for every term is far more expensive, and an inverted index does not work as-is with this architecture.

Back to the Roots

UniCOILContextual token scalars
Query
WhatisQdrant
BERT tokenizer → context
WhatisQdrant

BERT context: ht ∈ ℝ768 per token

One-layer NN · 768D → 1 scalar
Query token weights
What0.3is0.1Q1.5dra0.8nt0.9
Document 42
Qdrantisavectordatabase
BERT tokenizer → context
Qdrantisavectordatabase

BERT context: ht ∈ ℝ768 per token

One-layer NN · 768D → 1 scalar
Document token weights
Q1.7dra1.4nt0.8is0.6a0.2vector1.7database1.9
Exact token matching → scalar products → sum
is0.1 × 0.6
Q1.5 × 1.7
dra0.8 × 1.4
nt0.9 × 0.8
Σ matched products = 4.45
The query and document both get learned scalar weights. Use the maximum weight for repeated document tokens. Values are illustrative.
Universal COntextualized Inverted List (UniCOIL), made by the authors of COIL as a follow-up, goes back to producing a scalar value as the importance score rather than a vector, leaving unchanged all other COIL design decisions.
It optimizes resources consumption but the deep semantics understanding tied to COIL architecture is again lost.

Did we Solve the Vocabulary Mismatch Yet?

With the retrieval based on the exact matching, however sophisticated the methods to predict term importance are, we can’t match relevant documents which have no query terms in them. If you’re searching for “pizza” in a book of recipes, you won’t find “Margherita”.

A way to solve this problem is through the so-called document expansion. Let’s append words which could be in a potential query searching for this document. So, the “Margherita” document becomes “Margherita pizza”. Now, exact matching on “pizza” will work!

Document expansionBridge the vocabulary gap
Query
pizza
Document 42 · original
Margherita
No exact term match

pizza ≠ Margherita. The document cannot match this query through exact terms.

Query
pizza
Document 42 · expanded
Margheritapizza
Exact matching becomes possible

pizza now occurs in both query and document. The dashed term was added by document expansion.

Expansion adds potential query terms to a document. The query stays unchanged; the added pizza term creates an exact match.

There are two types of document expansion that are used in sparse neural retrieval: external (one model is responsible for expansion, another one for retrieval) and internal (all is done by a single model).

External Document Expansion

External document expansion uses a generative model (Mistral 7B, Chat-GPT, and Claude are all generative models, generating words based on the input text) to compose additions to documents before converting them to sparse representations and applying exact matching methods.

External Document Expansion with docT5query

docT5queryExternal expansion · queries
Example query · not model input

Can I send a money order from USPS as a business?

Document 42 · original passage

Sure you can. You can fill in whatever you want in the From section of a money order, so your business name and address would be fine. The price only includes the money order itself. You can hand deliver it yourself if you want, but if you want to mail it, you'll have to provide an envelope and a stamp.

docT5query · T5 generative model

Generate likely queries, one token at a time.

Append generated queries to the document
  • can you write a money order on a stub
  • can i mail money order to a contractor
  • how to send a money order
  • can you mail money order yourself
  • do i need a money order stamp
  • can you hand deliver a money order
  • how to send a money order without a stamp
  • can someone hand deliver money order
  • how long can someone deliver a money order

Repeated terms remain, so term frequencies can also increase.

Generated queries are appended to the passage before a separate retriever indexes it. Repeated words can raise term frequency.
docT5query is the most used document expansion model. It is based on the Text-to-Text Transfer Transformer (T5) model trained to generate top-k possible queries for which the given document would be an answer. These predicted short queries (up to ~50-60 words) can have repetitions in them, so it also contributes to the frequency of the terms if the term frequency is considered by the retriever.

The problem with docT5query expansion is a very long inference time, as with any generative model: it can generate only one token per run, and it spends a fair share of resources on it.

External Document Expansion with Term Independent Likelihood MODel (TILDE)

TILDEExternal expansion · terms
Example query · not model input

Can I send a money order from USPS as a business?

Document 42 · original passage

Sure you can. You can fill in whatever you want in the From section of a money order, so your business name and address would be fine. The price only includes the money order itself. You can hand deliver it yourself if you want, but if you want to mail it, you'll have to provide an envelope and a stamp.

TILDE · vocabulary likelihood model
1
Predict token probabilities in parallel

One distribution over the BERT vocabulary, conditioned on the passage.

2
Select the top-k terms

Append terms without repetitions; no query sentence generation.

Append predicted terms to the document
paygetorderscashneedcheckcostmuchsendlettercopytakeofficepaperreceivepostalandchargeorderingusesomeoneservicelongwaydepositpurchasehouseiteminstructionsdirectposttransfercarrydeliveryprinteditemsneedednumbercardpaidbuysentusputdonegoodsellcompanydocumentsfreerequiredbillcoform
TILDE appends independent terms without repetitions. A separate retriever then indexes the passage and added terms.

Term Independent Likelihood MODel (TILDE) is an external expansion method that reduces the passage expansion time compared to docT5query by 98%. It uses the assumption that words in texts are independent of each other (as if we were inserting in our speech words without paying attention to their order), which allows for the parallelisation of document expansion.

Instead of predicting queries, TILDE predicts the most likely terms to see next after reading a passage’s text (query likelihood paradigm). TILDE takes the probability distribution of all tokens in a BERT vocabulary based on the document’s text and appends top-k of them to the document without repetitions.

Problems of external document expansion: External document expansion might not be feasible in many production scenarios where there’s not enough time or compute to expand each and every document you want to store in a database and then additionally do all the calculations needed for retrievers. To solve this problem, a generation of models was developed which do everything in one go, expanding documents “internally”.

Internal Document Expansion

Let’s assume we don’t care about the context of query terms, so we can treat them as independent words that we combine in random order to get the result. Then, for each contextualized term in a document, we are free to pre-compute how this term affects every word in our vocabulary.

For each document, a vector of the vocabulary length is created. To fill this vector in, for each word in the vocabulary, it is checked if the influence of any document term on it is big enough to consider it. Otherwise, the vocabulary word’s score in a document vector will be zero. For example, by pre-computing vectors for the document “pizza Margherita” on a vocabulary of 50,000 most used English words, for this small document of two words, we will get a 50,000-dimensional vector of zeros, where non-zero values will be for a “pizza”, “pizzeria”, “flower”, “woman”, “girl”, “Margherita”, “cocktail” and “pizzaiolo”.

Sparse Neural Retriever with Internal Document Expansion

SPARTAInternal document expansion
Document 42
Qdrantisavectordatabase
BERT contextual encoding
Qdrantisavectordatabase

Context vector ht per document token

BERT vocabulary · 30,522 tokens

All tokens v, including terms absent from the document.

BERT embedding layer

Non-contextual vector ev per vocabulary token

Internal document expansion
1 · Dot product
ev · ht

Every v against every document token t

2 · Max over document tokens
mv = maxt ev · ht
3 · Threshold + log
dv = log(1 + ReLU(mv + bias))
Document · selected vocabulary slots
What0.8is0Q0dra0nt0.1
Query · unexpanded binary vector

What is Qdrant

What1is1Q1dra1nt1

30,522 slots; other slots: 0.

Select weights at query-token slots
What1 × 0.8
nt1 × 0.1
Σ selected weights = 0.9
The document expands across the vocabulary; the binary query bypasses expansion. Max similarity precedes the threshold and log transform. Selected weights and the score are illustrative.

The authors of the Sparse Transformer Matching (SPARTA) model use BERT’s model and BERT’s vocabulary (around 30,000 tokens). For each token in BERT vocabulary, they find the maximum dot product between it and contextualized tokens in a document and learn a threshold of a considerable (non-zero) effect. Then, at the inference time, the only thing to be done is to sum up all scores of query tokens in that document.

Why is SPARTA not a perfect solution? Trained on the MS MARCO dataset, many sparse neural retrievers, including SPARTA, show good results on MS MARCO test data, but when it comes to generalisation (working with other data), they could perform worse than BM25.

State-of-the-Art of Modern Sparse Neural Retrieval

SPLADE++Internal query + document expansion
Query
WhatisQdrant
BERT tokenizer → context
WhatisQdrant

ht ∈ ℝ768 per input token

Vocabulary head · 30,522 outputsLinear + GeLU + LayerNorm
xt,v = transformed(ht) · ev

ev: static BERT vocabulary embedding

Log activation → max pooling
at,v = log(1 + ReLU(xt,v + bv))
wv = maxinput tokens t at,v
Query · selected nonzero slots
vector0.5What0.8database0.2Q0.1dra0.3nt0.1
Document 42
Qdrantisavectordatabase
BERT tokenizer → context
Qdrantisavectordatabase

ht ∈ ℝ768 per input token

Vocabulary head · 30,522 outputsLinear + GeLU + LayerNorm
xt,v = transformed(ht) · ev

ev: static BERT vocabulary embedding

Log activation → max pooling
at,v = log(1 + ReLU(xt,v + bv))
wv = maxinput tokens t at,v
Document · selected nonzero slots
vector0.3What0.1database0.4nt0.5
Sparse query · sparse document → dot product
vector0.5 × 0.3
What0.8 × 0.1
database0.2 × 0.4
nt0.1 × 0.5
Σ shown named matches = 0.36
Both branches expand (dashed terms). Training uses sparsity regularization and distillation. Selected slots and their partial sum are illustrative; omitted slots are not scored here.
The authors of the Sparse Lexical and Expansion Model (SPLADE) family of models added dense model training tricks to the internal document expansion idea, which made the retrieval quality noticeably better.

  • The SPARTA model is not sparse enough by construction, so authors of the SPLADE family of models introduced explicit sparsity regularisation, preventing the model from producing too many non-zero values.
  • The SPARTA model mostly uses the BERT model as-is, without any additional neural network to capture the specificity of Information Retrieval problem, so SPLADE models introduce a trainable neural network on top of BERT with a specific architecture choice to make it perfectly fit the task.
  • SPLADE family of models, finally, uses knowledge distillation, which is learning from a bigger (and therefore much slower, not-so-fit for production tasks) model how to predict good representations.

One of the last versions of the SPLADE family of models is SPLADE++.
SPLADE++, opposed to SPARTA model, expands not only documents but also queries at inference time.

To try SPLADE++ yourself, follow the How to Generate Sparse Vectors with SPLADE tutorial: it walks through generating SPLADE++ vectors with FastEmbed and inspecting which terms the model expanded the text with.

SPLADE++ is also available as a paid model in Qdrant Cloud Inference.

Key Takeaways: When to Choose Sparse Neural Models for Retrieval

Sparse Neural Retrieval makes sense:

  • In areas where keyword matching is crucial but BM25 is insufficient for initial retrieval, semantic matching (e.g., synonyms, homonyms) adds significant value. This is especially true in fields such as medicine, academia, law, and e-commerce, where brand names and serial numbers play a critical role. Dense retrievers tend to return many false positives, while sparse neural retrieval helps narrow down these false positives.

  • Sparse neural retrieval can be a valuable option for scaling, especially when working with large datasets. It leverages exact matching using an inverted index, which can be fast depending on the nature of your data.

  • If you’re using traditional retrieval systems, sparse neural retrieval is compatible with them and helps bridge the semantic gap.

Was this page useful?

Thank you for your feedback! 🙏

We are sorry to hear that. 😔 You can edit this page on GitHub, or create a GitHub issue.