Reducing the p99 latency of a vector index
Why the tail degrades before the median: efSearch, metadata filtering, reindexing, and how memory pressure shows up in high percentiles.
Read the articleOur write-ups on vector search, RAG pipelines, index performance and the internal architecture of Vela.
What BM25 does well, what it cannot do, and why vector similarity search complements the lexical index rather than replacing it.
Read the articleMemory cost per vector, how dimension affects recall, quantization and truncation: how to decide before indexing a large corpus.
Read the articleWhy the tail degrades before the median: efSearch, metadata filtering, reindexing, and how memory pressure shows up in high percentiles.
Read the article