Original Reddit post

I’m currently looking at different vector stores for a project that I expect to scale, and the two things I actually care about are pretty simple: latency/p95 and recall. The problem is that almost every benchmark I find changes five things at once. Different dataset, different embedding model, different hardware, different index configuration, sometimes even different query setup. So I’ll see something like DB A doing 5ms and DB B doing 8ms, but I have no idea whether that 3ms difference would mean anything for my workload. Some use 1M vectors, some use synthetic data, some use proprietary datasets, and some outta nowhere have a very specific configuration that makes the comparison hard to reproduce. I’ve been looking more into larger open datasets like LanceDB and Qdrant-FineWeb-10B, the latter is the one I’ve been checking recently for the data and because LanceDB is calling it a day by getting to 10B by replicating the same dataset ~1,000 times. But if we keep the scale aside, I’m honestly more interested in the methodology than the numbers themselves. If you were evaluating Qdrant, Milvus, Weaviate, Elasticsearch, etc. for a real production workload, how would you actually set up the benchmark? Would you start with a public dataset and then move to your own data, or just benchmark your actual workload from the beginning? And do you care about a single latency/recall number, or is the recall-vs-latency curve a much better way to compare these systems? Also, how much weight would you actually put on benchmarks published by the vendors themselves? I feel like they can still be useful for understanding what a system is capable of, but I’m not sure I’d use them to make a production decision. submitted by /u/Pretty_Upstairs9035

Originally posted by u/Pretty_Upstairs9035 on r/ArtificialInteligence