Metadata Filtering: Pre-Filter vs Post-Filter Vector Search

LLM foundationsRetrieval and dataPublished Updated By Simon Budziak

Metadata filtering limits a search to records whose structured fields match set conditions. Fields such as tenant, document type, date, language or access level are stored beside each vector or document, and the filter decides which records may be ranked at all, before, during or after similarity scoring.

Elasticsearch applies a filter inside the knn clause during the approximate search, so k matching documents are returned, while a post-filter runs after the kNN step and can return fewer than k results even when enough matching documents exist (filtered kNN search, read 6 October 2026).

Why can a filtered vector search return too few results?

Approximate indexes such as HNSW in a vector database collect a fixed number of nearest candidates first. pgvector applies the filter after the index scan: with the default hnsw.ef_search of 40, a condition matching 10% of rows returns about 4 rows on average. Iterative index scans, added in version 0.8.0, keep scanning until enough rows match (pgvector README, version 0.8.7, read 6 October 2026). A strict filter on a plain approximate index can silently starve the result set. Qdrant takes another route: it adds HNSW graph edges for indexed payload values so the filter applies while the graph is searched, and falls back to the payload index for very strict filters. Those edges are built only after a payload index exists, so Qdrant recommends creating indexes before ingesting data (indexing docs, read 6 October 2026).

When should retrieval use metadata filters?

Filters help when semantic similarity alone cannot express a hard requirement. A support agent may search only the current customer’s records, or a policy assistant only approved documents. Hard access rules should constrain retrieval before results leave the data system, not rely on the model to ignore forbidden text.

What have we learned filtering our own knowledge base by metadata?

Our internal knowledge base is markdown files with YAML frontmatter, which exists so agents can select files by fields such as type, status and tags. A filter is only as good as the field it reads. Three rules came out of it:

Filters get tested like any other code path. The search tool in our open-source krs-mcp server filters by registry, entrepreneurs or associations and both by default, and its daily health check searches one registry and asserts that a known company comes back, so a filter that silently drops records fails the check.

What can go wrong with metadata filtering?

The needed fields must be attached during document chunking and kept current. Cloudflare Vectorize filters only on properties with a metadata index, allows up to 10 per index, and leaves out vectors upserted before that index existed (metadata filtering, updated 21 April 2026). Re-upsert existing records after adding a filterable field. Use hybrid search and filtering as separate controls: one improves relevance, while the other enforces scope for RAG and AI agent security.

Written with AI assistance and reviewed by Simon Budziak. The production notes come from systems Soba Labs builds and runs.

Frequently asked questions

What is metadata filtering in RAG?

It restricts which chunks a RAG retriever may return, using fields stored with each chunk such as source, date, language or permission group, so the model only sees text from the allowed scope.

What is the difference between pre-filtering and post-filtering?

Pre-filtering limits eligible records before or during vector search. Post-filtering retrieves nearest vectors first and then removes records that fail the filter.

Can metadata filtering enforce document permissions?

It can contribute to enforcement when permission metadata is current and the search system applies it before returning results. The application must still authenticate the requester.

Summarize this page with

Train your team to build this