Elasticsearch applies a filter inside the knn clause during the approximate search, so k matching documents are returned, while a post-filter runs after the kNN step and can return fewer than k results even when enough matching documents exist (filtered kNN search, read 6 October 2026).
Why can a filtered vector search return too few results?
Approximate indexes such as HNSW in a vector database collect a fixed number of nearest candidates first. pgvector applies the filter after the index scan: with the default hnsw.ef_search of 40, a condition matching 10% of rows returns about 4 rows on average. Iterative index scans, added in version 0.8.0, keep scanning until enough rows match (pgvector README, version 0.8.7, read 6 October 2026). A strict filter on a plain approximate index can silently starve the result set. Qdrant takes another route: it adds HNSW graph edges for indexed payload values so the filter applies while the graph is searched, and falls back to the payload index for very strict filters. Those edges are built only after a payload index exists, so Qdrant recommends creating indexes before ingesting data (indexing docs, read 6 October 2026).
When should retrieval use metadata filters?
Filters help when semantic similarity alone cannot express a hard requirement. A support agent may search only the current customer’s records, or a policy assistant only approved documents. Hard access rules should constrain retrieval before results leave the data system, not rely on the model to ignore forbidden text.
What have we learned filtering our own knowledge base by metadata?
Our internal knowledge base is markdown files with YAML frontmatter, which exists so agents can select files by fields such as type, status and tags. A filter is only as good as the field it reads. Three rules came out of it:
- Categorical fields use short, controlled vocabularies, so a filter matches one stable token. A linter before every commit rejects a file with no type, a type outside the fixed list, or a status outside its documented values.
- The header changes in the same edit as the body. Our most repeated defect was a record whose body was updated while its header was not, so it read current and filtered stale. Two records ran 8 and 4 days stale, each header pointing the opposite way from the truth.
- Values are checked with the real parser. In YAML, an unquoted
#after a space starts a comment, so the value is silently cut short and a filter reads back the truncated value as fact.
Filters get tested like any other code path. The search tool in our open-source krs-mcp server filters by registry, entrepreneurs or associations and both by default, and its daily health check searches one registry and asserts that a known company comes back, so a filter that silently drops records fails the check.
What can go wrong with metadata filtering?
The needed fields must be attached during document chunking and kept current. Cloudflare Vectorize filters only on properties with a metadata index, allows up to 10 per index, and leaves out vectors upserted before that index existed (metadata filtering, updated 21 April 2026). Re-upsert existing records after adding a filterable field. Use hybrid search and filtering as separate controls: one improves relevance, while the other enforces scope for RAG and AI agent security.