Qdrant vs Weaviate vs Milvus vs pgvector: which vector database should you use?

7 minUpdated:
Qdrant vs Weaviate vs Milvus vs pgvector: which vector database should you use?

Use pgvector if you already run Postgres and your dataset is moderate; Qdrant for a lean, fast dedicated vector store with strong filtering; Weaviate if you want built-in vectorisation modules and hybrid search; Milvus when you expect very large collections and can operate a distributed system.

What are these four options?

Qdrant, Weaviate and Milvus are dedicated open-source vector databases. pgvector is a PostgreSQL extension that adds a vector column type and similarity search to the database you may already run.

All four store embeddings and return nearest neighbours, usually through approximate indexes such as HNSW. The differences are in architecture, operations, filtering, and how much surrounding machinery they ship.

How do they compare side by side?

QdrantWeaviateMilvuspgvector
LicenceApache 2.0BSD-3-ClauseApache 2.0PostgreSQL licence
LanguageRustGoGo and C++C (Postgres extension)
DeploymentSingle binary or cluster; managed cloud offeredSingle node or cluster; managed cloud offeredStandalone or distributed with several components; managed cloud offeredAnywhere Postgres runs, including most managed Postgres services
FilteringRich payload filters combined with vector searchFilters plus hybrid keyword and vector searchScalar filtering and multiple index typesFull SQL, joins and transactions
ExtrasQuantisation options, snapshotsModules that call embedding and generation modelsMany index types including GPU optionsEverything Postgres gives you
Best forDedicated RAG store with heavy filteringTeams wanting batteries-included semantic searchVery large scale collectionsApps already on Postgres
Trade-offAnother service to runMore concepts and configurationHeaviest to operateTuning and scale limits of one Postgres instance

When is pgvector enough?

For many products pgvector is the right first answer. Your embeddings sit next to users, documents and permissions, so a single SQL query can filter by tenant, join metadata and rank by similarity inside one transaction.

You avoid syncing two systems, which removes a whole class of bugs where the vector store and the source of truth drift apart. Backups, access control and migrations stay on tooling you already know.

It starts to strain when vector workloads compete with transactional traffic, when index builds on large tables become slow, or when you need features such as built-in quantisation strategies and horizontal sharding of vectors.

Why pick Qdrant?

Qdrant is a focused vector engine written in Rust. It is easy to start as a single container, and its payload filtering is designed to work together with the vector index rather than as a slow post-filter.

That makes it a strong fit for RAG over documents with many attributes: tenant, language, date, access level. Its API is compact and client libraries are available for the common languages.

The trade-off is operational: it is one more stateful service with its own backups, upgrades and monitoring.

Qdrant also supports sparse vectors alongside dense ones, which opens the door to hybrid retrieval, and several quantisation modes that trade a little recall for much lower memory use. Those features matter once collections grow beyond what fits cheaply in RAM.

Teams that like it tend to cite predictable behaviour and a small operational footprint. Teams that leave it usually do so because they wanted their vectors inside the main database after all.

Why pick Weaviate or Milvus?

Weaviate leans toward being a complete semantic search platform. Its modules can call embedding models at write and query time, and hybrid search blends keyword and vector relevance, which helps when exact terms such as product codes matter.

Milvus is built for scale. Its distributed mode separates storage, indexing and query roles, and it offers a broad choice of index types. That flexibility suits very large collections but brings more moving parts to run.

Both have lighter modes for development. Test your real workload on the production-style deployment, not only on the laptop version.

A useful way to frame the choice: Weaviate saves application code by doing more inside the database, while Milvus saves hardware headaches at large scale by distributing work. If neither of those problems is yours yet, a simpler option probably wins.

How much does each cost to run?

The software itself is free to self-host in all four cases, so cost comes from memory, storage and people. Vector indexes such as HNSW perform best when held in RAM, and RAM is usually the biggest line item as collections grow.

Embedding dimensions matter directly: doubling dimensions roughly doubles raw vector storage. Qdrant, Weaviate and Milvus offer quantisation options that shrink memory use at some cost in recall, and pgvector supports reduced-precision vector types in recent versions; check the documentation for your version.

The hidden cost is operations. pgvector adds almost nothing if you already run Postgres well. A single Qdrant or Weaviate node is a modest extra service. A distributed Milvus deployment brings several components and dependencies that someone must monitor, upgrade and back up.

Managed clouds from each vendor trade that effort for a bill. Compare them against your own time honestly, and check current pricing pages rather than relying on old blog posts.

How does filtering change the decision?

Real RAG queries are rarely “find the nearest vectors”. They are “find the nearest vectors this user may see, in this language, updated this year”. How a database combines filters with approximate search decides whether you get fast, complete results.

Naive post-filtering retrieves the top results first and then discards those that fail the filter. With selective filters you can end up with too few results or none at all. Dedicated engines such as Qdrant and Weaviate invest heavily in filtering that works during the index traversal.

In Postgres, the planner decides how to combine a vector index with other conditions, and results depend on your indexes and settings. Recent pgvector releases improved filtered search, but you should test selective filters explicitly with realistic data.

Multi-tenant SaaS is the classic case. Decide early whether each tenant gets its own collection, a partition, or a shared collection with a tenant filter, because that choice shapes performance and deletion workflows.

Which should you choose?

To choose without regret, start by estimating how many vectors you will hold in year one and at what dimensions. Memory, not query speed, is usually the first limit you hit.

Next, write down your real filters. If most queries filter by tenant or access level, test filtered search specifically, with the selectivity you expect in production.

Then benchmark recall and latency on your own embeddings, not on public datasets, because data distribution changes results. Hide the store behind a small repository interface so a later migration touches one module.

Finally, decide who owns backups, upgrades and monitoring before you add a new stateful service. If nobody does, that is a strong argument for pgvector or a managed offering.

  • Already on Postgres, dataset in a range one database handles comfortably: pgvector.
  • Dedicated RAG backend with complex metadata filters and a small ops team: Qdrant.
  • Semantic search product where hybrid keyword and vector ranking and built-in vectorisation save work: Weaviate.
  • Very large embedding collections and a team comfortable running distributed infrastructure: Milvus.
  • Prototype or hackathon: pgvector or single-node Qdrant, whichever your stack makes easier.
  • Strict multi-tenancy with row-level security already in Postgres: pgvector keeps permissions in one place.

Common mistakes

RepoLoot’s catalog lists RAG starters built on each of these stores, which is a quick way to see real integration code before you commit.

  • Choosing for scale you will not reach and paying the operational cost from day one.
  • Forgetting that re-embedding with a new model means rebuilding the whole index.
  • Measuring only latency and ignoring recall, so fast results are quietly worse.
  • Post-filtering results in application code, which returns too few hits when filters are selective.
  • Keeping no copy of the source text, making re-indexing impossible without the original pipeline.

Frequently asked questions

Is pgvector fast enough for production?
For many production apps, yes, especially with an HNSW index and a dataset that fits comfortably in memory. It becomes a harder fit when vector search load competes with transactional load or collections grow very large. Test with realistic data and concurrency before deciding.
Can I self-host all four for free?
Yes. Qdrant, Weaviate and Milvus publish open-source servers you can run yourself, and pgvector is an open-source Postgres extension. Each vendor also sells a managed cloud; check current licence files if you plan to offer the database itself as a service.
Which is best for hybrid search?
Weaviate has hybrid keyword and vector search as a core feature. Qdrant supports sparse vectors that enable hybrid approaches, and in Postgres you can combine pgvector with full-text search in SQL. The best choice depends on how much ranking logic you want to own.
How hard is it to switch later?
Moving vectors is easy; moving behaviour is harder. Filters, hybrid scoring and index settings differ, so results change. Keep original text and metadata, isolate the store behind an interface, and treat a migration as a re-index plus an evaluation run.
Free for builders

Get a hand-picked shortlist of repos for your project

Tell us what you are building. A person — not a bot — reviews it and replies within 48 hours with the catalog projects that fit, including licence and difficulty notes.

We use your email only for this request. Privacy policy

Related guides