Sparse Retrieval Explained: BM25, SPLADE and Exact Matches

LLM foundationsRetrieval and dataPublished Updated By Simon Budziak

Sparse retrieval is search that ranks documents by the exact terms they share with a query. Each document becomes a vector that is zero for almost every word in the vocabulary. BM25 weights the few matching terms by rarity and frequency, which makes sparse retrieval strong for names, codes, identifiers and rare phrases.

Elasticsearch uses BM25 as its default similarity, with two tuning parameters: k1 (default 1.2) sets how quickly extra occurrences of a term stop adding score, and b (default 0.75) sets how much document length discounts term counts (similarity settings, read 6 October 2026).

How does sparse retrieval work?

The index is an inverted index: for every term, a list of the documents that contain it. At query time only the lists for the query’s terms are scored, so a rare term such as an error code counts far more than a common word. Exact lexical matching is its advantage, not an outdated compromise. Its blind spot is wording: a passage that says the same thing in other words shares no terms and never scores. Reranking can inspect the top candidates with a stronger model, while a vector database serves the dense path built on embeddings.

What is a sparse retriever, and how is SPLADE different from BM25?

A sparse retriever turns a query into term weights and returns the best matching documents from an inverted index. BM25 weights only the words present. Learned sparse retrievers use a neural model to choose and weight terms, including related terms the text never used. SPLADE (Formal, Piwowarski and Clinchant, July 2021) adds sparsity regularization so most weights stay zero, keeping inverted index retrieval efficient with results competitive with dense and sparse baselines. Learned sparse vectors still run on ordinary search infrastructure: Elasticsearch stores them in a sparse_vector field used by its ELSER model (field reference, read 6 October 2026).

What have we learned relying on keyword search in our own agent systems?

Our agents’ working knowledge is plain markdown in git, found by keyword search and file paths. We chose to keep it there rather than move agent memory into a vector or graph memory framework. Keyword search fails quietly, so we treat a short or empty result as unproven. Three rules came out of that:

Should RAG use sparse or dense retrieval?

Evaluate both on real questions. Semantic search wins when wording varies. Sparse retrieval protects identifiers and short precise queries. Hybrid search combines both rankings. It is often simpler for mixed enterprise data than forcing one method to do every job. OpenClaw’s built-in agent memory, for one, pairs FTS5 keyword search with BM25 scoring, vector search and a hybrid mode (memory docs, read 6 October 2026), as covered in our Hermes Agent memory post.

Written with AI assistance and reviewed by Simon Budziak. The production notes come from systems Soba Labs builds and runs.

Frequently asked questions

Is sparse retrieval just keyword search?

Classical sparse retrieval, including BM25, is keyword search with term weighting. Learned sparse methods such as SPLADE use a neural model to weight and expand terms while keeping a sparse representation.

When is sparse retrieval better than vector search?

It is often better for identifiers, product names, legal phrases, error codes, and rare terms that should match exactly rather than by general semantic similarity.

What is the difference between sparse and dense retrieval?

Sparse retrieval stores weights for vocabulary terms and matches documents that share words with the query. Dense retrieval stores an embedding in which every value is used and matches on meaning, so it finds paraphrases that share no words.

Summarize this page with

Train your team to build this