Elasticsearch uses BM25 as its default similarity, with two tuning parameters: k1 (default 1.2) sets how quickly extra occurrences of a term stop adding score, and b (default 0.75) sets how much document length discounts term counts (similarity settings, read 6 October 2026).
How does sparse retrieval work?
The index is an inverted index: for every term, a list of the documents that contain it. At query time only the lists for the query’s terms are scored, so a rare term such as an error code counts far more than a common word. Exact lexical matching is its advantage, not an outdated compromise. Its blind spot is wording: a passage that says the same thing in other words shares no terms and never scores. Reranking can inspect the top candidates with a stronger model, while a vector database serves the dense path built on embeddings.
What is a sparse retriever, and how is SPLADE different from BM25?
A sparse retriever turns a query into term weights and returns the best matching documents from an inverted index. BM25 weights only the words present. Learned sparse retrievers use a neural model to choose and weight terms, including related terms the text never used. SPLADE (Formal, Piwowarski and Clinchant, July 2021) adds sparsity regularization so most weights stay zero, keeping inverted index retrieval efficient with results competitive with dense and sparse baselines. Learned sparse vectors still run on ordinary search infrastructure: Elasticsearch stores them in a sparse_vector field used by its ELSER model (field reference, read 6 October 2026).
What have we learned relying on keyword search in our own agent systems?
Our agents’ working knowledge is plain markdown in git, found by keyword search and file paths. We chose to keep it there rather than move agent memory into a vector or graph memory framework. Keyword search fails quietly, so we treat a short or empty result as unproven. Three rules came out of that:
- A word you chose is an assumption. One keyword sweep returned a clean list of 17 documents out of 20; the other 3 never used the searched word. To list a collection, we query the container, not a term.
- The tool can skip files silently. A ripgrep-style search honors
.gitignore, and our nested repositories are ignored by design, so a search from the top saw only the outer one. A non-UTF-8 file is treated as binary and skipped: one count came back empty and would have been reported as 0 instead of 15. Every absence check now includes a term that must return hits. - Identifiers get exact matching. In krs-mcp, our open-source server for the Polish company register, a company name goes to the search query, while KRS, NIP and REGON numbers are separate exact fields.
Should RAG use sparse or dense retrieval?
Evaluate both on real questions. Semantic search wins when wording varies. Sparse retrieval protects identifiers and short precise queries. Hybrid search combines both rankings. It is often simpler for mixed enterprise data than forcing one method to do every job. OpenClaw’s built-in agent memory, for one, pairs FTS5 keyword search with BM25 scoring, vector search and a hybrid mode (memory docs, read 6 October 2026), as covered in our Hermes Agent memory post.