Quick Answer: No single vector database for RAG wins: pick Pinecone for managed, Qdrant, Weaviate or Milvus to self-host, and pgvector if you already run Postgres. Chroma is another Apache 2.0 option with a managed cloud. Decide on deployment model, hybrid search with filtering, and vector count, then test two candidates on your own queries.
Every engine in this guide stores embeddings, runs approximate nearest-neighbor search and filters on metadata. So the vector store is rarely what makes a RAG answer right or wrong. Chunking, the embedding model, hybrid keyword matching and permission filters decide that far more often.
What does differ is where the database runs, under which licence, how it mixes keyword and vector scores, and how much infrastructure your team has to operate. That's what this comparison covers, using each vendor's own documentation as read on 25 September 2026.
Which vector database is best for RAG?
There is no universal best vector database for RAG. The right pick follows three facts about your system: where it must run, whether it needs hybrid keyword-plus-vector search with filters, and how many vectors it will hold.
Here is the shortlist, grouped by the situation each one fits:
- Managed, zero-ops: Pinecone. You get an API and no servers to run. Self-hosting is limited to its Enterprise bring-your-own-cloud option.
- Open source, filter-heavy RAG: Qdrant. Apache 2.0, indexed payload filtering and dense-plus-sparse hybrid queries.
- Open source, hybrid built in: Weaviate. It fuses BM25F keyword results with vector results in one query.
- Open source, very large corpora: Milvus. Its distributed mode is documented for 100 million to tens of billions of vectors, at the cost of more moving parts.
- Lightweight and quick to start: Chroma. It runs locally or self-hosted, with a managed cloud when you outgrow a laptop.
- Already on PostgreSQL: pgvector. Vectors live next to your relational data, with SQL filters and joins.
If you're after the best vector database for a first production RAG system, start from your deployment constraint, not from a benchmark chart. The deployment constraint usually shortens the list in one step.
How do the leading vector databases compare for RAG?
The six options split cleanly on licence and hosting, and much less on search features. A vector database comparison that only lists "supports filtering" will show Yes in every row, so the notes column carries the real differences.
| Database | Open source | Self-host in your environment | Managed cloud | Hybrid search | Metadata filtering | Notes |
|---|---|---|---|---|---|---|
| Pinecone | Not documented | Enterprise tier | Yes | Yes | Yes | Offered as a managed service; BYOC runs in your AWS, GCP or Azure account on the Enterprise plan |
| Qdrant | Yes | Yes | Yes | Yes | Yes | Apache 2.0; Hybrid Cloud and Private Cloud are sold through Qdrant's sales team |
| Weaviate | Yes | Yes | Yes | Yes | Yes | BSD-3-Clause; the licence file reserves a separate, key-gated Weaviate License for code in the repo's wl directory |
| Milvus | Yes | Yes | Yes | Yes | Yes | Apache 2.0; Lite, Standalone and Distributed modes; managed through Zilliz Cloud |
| Chroma | Yes | Yes | Yes | Chroma Cloud only | Yes | Apache 2.0; self-hosted Chroma filters on document text, but hybrid ranking with rank fusion (the Search API) is Chroma Cloud only |
| pgvector | Yes | Yes | Yes | Yes | Yes | Postgres extension under the same licence text as PostgreSQL; hybrid means pairing it with Postgres full-text search; managed through hosted Postgres providers |
Capabilities as documented by each vendor on 25 September 2026; links in the text.
Three cells need a closer read. Pinecone's self-hosting is its bring-your-own-cloud deployment in your cloud account, not software you install on your own servers. Chroma's hybrid ranking runs only in Chroma Cloud today. And "managed cloud" for Milvus and pgvector means third-party hosts: Zilliz for Milvus, and the hosted Postgres providers the pgvector project lists.
Do you need a dedicated vector database, or is pgvector enough?
For many RAG systems, pgvector is enough. If your application already runs on PostgreSQL and the corpus is in the low millions of chunks, one database is simpler to secure, back up and audit than two.
What pgvector gives you
The pgvector extension adds a vector type, HNSW and IVFFlat indexes, and similarity operators to Postgres. Its README documents indexing up to 2,000 dimensions for the standard vector type and 4,000 with half precision. Filtering is a normal SQL WHERE clause, and hybrid search pairs it with Postgres full-text search.
When a dedicated engine earns its place
A dedicated engine starts to pay off when vector search load would compete with your transactional workload. It also helps when you need sparse vectors and rank fusion as first-class features, or when the corpus heads toward hundreds of millions of vectors. Choose pgvector when your data and permissions already live in Postgres and you'd rather not run another stateful service.
Whether you buy a one-click knowledge base tool or build a custom RAG pipeline, the vector store is a component, not the decision. If you're still weighing RAG as a service against your own pipeline, settle that first. Then benchmark pgvector against one dedicated engine on your own corpus and filters, because generic ANN benchmarks rarely include filtered queries.
Which vector databases run on-premise or in a private cloud?
Five of the six can run entirely on infrastructure you control: Qdrant, Weaviate, Milvus, Chroma and pgvector. Pinecone is the exception, since its closest option is running inside your own cloud account.
Open-source engines you deploy yourself
Qdrant, Milvus and Chroma ship under Apache 2.0, and pgvector installs into any Postgres you run. Qdrant also offers Hybrid Cloud, which runs managed clusters on your own infrastructure, and Private Cloud, an isolated deployment its pricing page lists for air-gapped setups. Both go through its sales team. The Milvus project documents Lite, Standalone and a Kubernetes-native distributed mode. Weaviate runs on Docker or Kubernetes. Its code is BSD-3-Clause, but its licence file reserves a separate Weaviate License, enabled only by a license key, for code in the repository's wl directory.
Pinecone inside your cloud account
Pinecone's bring-your-own-cloud deployment runs on AWS, GCP or Azure and requires its Enterprise plan. Per its documentation, vectors, metadata and queries stay in your cloud account. That can satisfy a data-residency review, but it isn't an on-premise or air-gapped deployment.
What an air-gapped RAG stack also needs
For an air-gapped deployment, the vector database is the easy part. The embedding model, the reranker and the LLM must also run inside the boundary, and any licensed feature must work without calling home. Ask each vendor that question directly before you design around it. The source documents also need a home upstream, and data lake vs data warehouse for AI workloads covers where they should live.
How do vector databases compare on filtering and hybrid search?
All six filter on metadata, and all six document hybrid search, though Chroma offers it only in Chroma Cloud. The differences are in how keyword and vector scores are combined and where filters apply.
Why hybrid search matters for enterprise documents
Embeddings are weak at exact tokens: part numbers, error codes, contract clause IDs, people's names. Hybrid search adds a keyword signal so those still match. Weaviate's hybrid search fuses BM25F and vector results, with relative score fusion as the default since v1.24. Qdrant's hybrid queries combine dense and sparse vectors and fuse them with Reciprocal Rank Fusion.
Pinecone's documentation combines a keyword signal with a semantic signal and treats metadata filtering as a separate lever that narrows candidates before ranking. Milvus supports BM25 full-text search alongside dense vectors in the same collection. Chroma's hybrid ranking with rank fusion lives in its Search API, which its documentation says is available in Chroma Cloud only; self-hosted Chroma can still filter on document text. pgvector relies on Postgres full-text search for the keyword half.
Why filters decide whether answers are allowed
Filters carry permissions. If a user may only see HR documents from their region, that condition has to apply inside the vector query, not after it. Post-filtering a top-10 result can return three allowed chunks, or none. Qdrant, for example, recommends payload indexes on every field you filter on.
The tradeoff across these engines is mostly scale, filtering behavior, hybrid mechanics, running cost and operational burden. Weaviate suits teams that want hybrid search built into one engine and can handle a richer configuration. Milvus suits teams planning for very large vector counts that can run a distributed system.
How do you test retrieval quality before you commit?
Build an evaluation set from your own traffic and run it against two candidate stores with everything else held fixed. A shortlist of two, measured on your data, beats any public leaderboard.
A workable test looks like this:
- Collect 100 to 200 real questions, each labelled with the passages that answer it.
- Fix the chunking strategy and the embedding model so only the store changes.
- Measure recall@k (did the right chunk appear in the top k?) and MRR (how high did it rank?).
- Run every query with the metadata and permission filters production will use.
- Record p95 latency under those filters, not on unfiltered queries.
- Repeat after a bulk update, because index freshness affects recall.
Treat any published vector database comparison as a shortlist, not a verdict. Public benchmarks such as BEIR measure embedding models across generic datasets. They can't tell you how your chunks, your filters and your users' wording behave together.
What mistakes should you avoid when choosing a vector database?
Most failed choices of a vector database for RAG come from testing the wrong thing. These are the ones that cost teams the most rework:
- Choosing on benchmark QPS. Unfiltered throughput says little about a RAG workload that filters every query by tenant, role or date.
- Leaving permissions for later. Retrofitting access control into metadata means re-indexing, so design the filter fields on day one.
- No re-indexing plan. Changing the embedding model means re-embedding the whole corpus. Plan how you'll run old and new indexes side by side.
- Mismatched embedding dimensions. Check the model's output size against each engine's index limits before you commit.
- Expecting the store to fix weak chunks. Poor chunking produces poor retrieval in every database on this list.
- Skipping the licence review. Read the licence file for the exact version you'll self-host, especially for mixed-licence projects.
How Origins AI runs retrieval in Velocity AI Suite
Origins AI (originshq.com) is an AI-augmented engineering company that builds and deploys self-hosted enterprise AI. Its knowledge and retrieval product, Origins AI Velocity AI Suite, treats the vector store as your choice rather than its own. The product page lists compatibility with Pinecone, Chroma, Weaviate and proprietary options, plus SQL, Postgres and MongoDB sources.
The suite's Knowledge Foundation layer handles data intake, document intelligence and retrieval indexing. Origins AI reports support for 1,900+ data sources and 91+ document formats. Models are bring-your-own, including OpenAI, Anthropic and open-source options. According to its products hub, every product can be deployed on-premise, in your private cloud account or air-gapped, with its team handling implementation and integration.
In on-premise and air-gapped modes, documents, embeddings and queries stay inside your network. A hybrid deployment sends retrieved context to the hosted model you pick. The listed security controls are encryption at rest and in transit, secure authentication, continuous monitoring and least-privilege access.
The same indexed knowledge base can serve internal chat, customer support and embedded assistants, which is the use case compared in our guide to knowledge base builders for chat and support. If you only need a vector index for one existing app, pgvector or a managed Pinecone index is less to run, and that's the better fit.
Talk to an engineer
Choosing a vector store for a self-hosted RAG system, or deciding whether you need one at all? Book a call with an Origins AI engineer to walk through your deployment constraints and retrieval test plan.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


