Lists / AI infra
Best vector databases for AI apps (2026): What ChatGPT, Claude & Gemini Recommend
pgvector, Qdrant and Pinecone lead 50 recorded AI answers on vector databases for RAG. A ranked shortlist with who should pick each, the per-model split and pricing.
pgvector, named in 49 of 50 recorded AI answers and first in 16 of 50, is the vector database the models name most for RAG. Qdrant matches its reach. Pinecone opens more answers, first in 19 of 50. Those three are the shortlist, and the right pick depends on the stack you already run.
This page counts which products 10 AI models name when asked about vector databases for RAG, in the 2026-09 edition. It does not test the products.
TL;DR
- pgvector is the measured leader for every model, but the answers mostly attach a condition: a team that already runs Postgres.
- Pinecone opens more answers than any rival and is framed as the managed, no-operations pick.
- Qdrant ties pgvector on reach and is the default the GPT-5.6 answers reach for, citing Qdrant’s own documentation.
- Weaviate and Milvus sit in 45 answers each and open none. The models file them as specialists for hybrid search and for scale.
- Choose by stack: Postgres already running points to pgvector, zero operations to Pinecone, filter-heavy open source to Qdrant.
1. pgvector
Pick pgvector if your application already runs on PostgreSQL and you want embeddings stored next to the data they describe.
Measured: named 49 of 50 (pgvector 98%), first 16 of 50 (pgvector 32%), average position 2.82.
pgvector is the measured leader for all ten models. Its only miss is a Sonar Reasoning Pro response that stops mid-sentence before it names any product. The answers that recommend it usually attach a condition, a team that already runs Postgres, and the section on the first slot below explains why its lead is softer than the row suggests.
The ZenML guide explains the appeal. If an application already uses Postgres, pgvector keeps data and vectors in one database, with full SQL querying power on the same data that backs the app. The same guide warns that Postgres isn’t designed for massive vector workloads.
That’s the trade a buyer accepts for running one database instead of two.
Pros
- Named in every answer from nine of the ten models
- Adds vector columns to an existing Postgres database, so there is no second system to sync
- Free as software, with Postgres hosting as the only cost
Cons
- First in 16 of 50, behind Pinecone’s 19
- Lags specialised vector engines on high-dimensional data, per the ZenML guide
- Large embedding tables can slow down other database operations
Pricing: free open-source extension. Costs are for PostgreSQL hosting. Best for: teams already on PostgreSQL that want to add RAG without a new database
2. Qdrant
Pick Qdrant if retrieval is the core of the product and results must be narrowed by tenant, permission or date before the final chunks are returned.
Measured: named 49 of 50 (Qdrant 98%), first 12 of 50 (Qdrant 24%), average position 2.24.
Qdrant ties pgvector on reach and misses the same truncated answer. The difference is who makes it the default. The GPT-5.6 answers most often do, and when they do they cite Qdrant’s own documentation. The host qdrant.tech is the most-cited vendor-owned host in the run, at 33 citations. Its average position is earlier than pgvector’s, so the answers tend to reach it sooner even when they open with something else.
Pros
- Average position 2.24, second only to Pinecone
- Payload indexes speed up filtered vector queries
- Runs locally, in Docker, on Kubernetes or as Qdrant Cloud
- Built-in quantization cuts memory use and cost
Cons
- Sharding is newer and less battle-tested than in older databases, per the ZenML guide
- Needs a separate embedding model or service before vectors are stored
Pricing: open source. Qdrant Cloud offers a free 1GB cluster, then usage-based managed clusters. Best for: filter-heavy retrieval
3. Pinecone
Pick Pinecone if nobody on the team should be running database infrastructure.
Measured: named 47 of 50 (Pinecone 94%), first 19 of 50 (Pinecone 38%), average position 2.04.
Pinecone trails the top two on reach and leads them on position. The gap on reach comes mostly from the OpenAI family (Pinecone 86.7% of its 15 answers). The answers that lead with it treat it as the managed pick.
Braintrust’s guide describes the same role. Pinecone runs serverless indexes as its default architecture, and its bring-your-own-cloud option can run the data plane in the buyer’s own cloud account. The guide frames managed hosting as scaling without owning upgrades, backups, monitoring and capacity planning.
For a team without database operators, that’s the whole case.
Pros
- First in 19 of 50 answers, the most of any product
- Earliest average position in the table, at 2.04
- Indexing, scaling, backups, filtering and hybrid search all run as a managed service
- Real-time indexing keeps new data searchable without batch re-indexing
- Every Anthropic and Google answer names it (Pinecone 100%)
Cons
- Managed only, with no general self-hosted version
- Usage-based billing means costs rise with queries, writes and storage
- GPT-5.6 Sol, GPT-5.6 Terra and Sonar Reasoning Pro each name it in 4 of 5
Pricing: free Starter tier. Standard carries a $50 monthly minimum and Enterprise a $500 monthly minimum. Best for: managed RAG with low infrastructure work
4. Weaviate
Pick Weaviate if hybrid keyword-and-vector search is central to the application and the team wants an open-source base.
Measured: named 45 of 50 (Weaviate 90%), first 0 of 50 (Weaviate 0%), average position 4.02.
Weaviate appears in 45 answers and opens none of them. Its 90-point gap between named and first is the widest in the table, shared only with Milvus. The models treat it as a specialist. They reach for it once the answer turns to hybrid search or multi-tenant setups, not as an opening default.
Braintrust’s guide gives the reason. Weaviate combines BM25 keyword search with vector search and lets teams tune how the two result sets are fused. Hybrid search helps with exact strings such as product names, error codes and ticket IDs.
A buyer whose queries depend on strings like those has the clearest reason to test Weaviate.
Pros
- Six of the ten models name it in all five of their answers
- Modules cover vectorization, reranking and generative workflows
- BSD-3 licensed, so self-hosting carries no licence fee
Cons
- Never first in 50 answers
- Schema design is critical, and vectorizer module costs can add up
- Self-hosting still needs configuration, scaling, monitoring and tuning
Pricing: open source. Weaviate Cloud serverless plans start at $25 a month. Best for: open-source RAG with hybrid search
5. Milvus
Pick Milvus if the dataset is very large and the team is prepared to run distributed vector search.
Measured: named 45 of 50 (Milvus 90%), first 0 of 50 (Milvus 0%), average position 4.64.
Milvus matches Weaviate on reach and on the empty first column, with a later average position. The answers file it under scale, as the option for billion-vector or distributed deployments once the default no longer fits. Its soft spot is Perplexity. Sonar Reasoning Pro names it in 3 of 5 answers.
The ZenML guide gives the reason for the scale role. Milvus separates compute and storage in a distributed architecture. It is a graduate project of the LF AI & Data Foundation. IBM’s RAG explainer names it beside Pinecone as a platform teams build with.
Pros
- Scales to billions or even trillions of vectors, per the ZenML guide
- HNSW, IVF, PQ and CAGRA indexes let teams trade recall against speed
- Anthropic and Google families cover it fully (Milvus 100%)
Cons
- Complex to operate, with clustering and resource settings that need careful tuning
- Latest average position of the top five, at 4.64
- Perplexity family coverage sits at Milvus 70%, the lowest family figure for any top-five product
Pricing: open source. Zilliz Cloud has a free tier, Dedicated plans from $99 a month and Serverless at $0.30 per GB per month. Best for: very large collections and distributed deployments
6. Chroma
Pick Chroma to get a prototype or small RAG app retrieving quickly, before committing to production infrastructure.
Measured: named 24 of 50 (Chroma 48%), first 2 of 50 (Chroma 4%), average position 6.42.
Chroma carries the sharpest model split in the category. Claude Fable 5 names it in 4 of 5 answers. GPT-5.6 Terra, Gemini 3.6 Flash, Sonar Pro and Sonar Reasoning Pro name it once each. A buyer’s chance of hearing about Chroma depends mostly on which assistant they ask. The Claude answers place it in the early-stage slot, paired with pgvector as the move-fast option for small projects. Its two first mentions are the only ones outside the top three.
Braintrust’s guide calls Chroma the lightweight option for local development, prototypes and smaller applications that need vector search without a separate database service.
Pros
- Starts inside a Python application, with Chroma Cloud as the hosted path
- Claude Opus 5 and Claude Sonnet 5 name it in all five answers
Cons
- GPT-5.6 Sol never names it
- Hybrid search is limited, strongest in the Chroma Cloud Search API
- Teams needing strict high availability or large-scale tenant isolation may later move to a heavier database
Pricing: no public pricing is recorded. Best for: prototypes and smaller apps
7. Elasticsearch
Pick Elasticsearch if the organisation already runs it and wants vector search without a new platform.
Measured: named 23 of 50 (Elasticsearch 46%), first 0 of 50 (Elasticsearch 0%), average position 6.65.
Elasticsearch appears in almost half the answers but thinly for most models. GPT-5.6 Sol and Claude Opus 5 name it in 4 of 5. Claude Sonnet 5, Claude Fable 5 and Sonar Pro name it once each. Where the answers explain it, they present it as the option for teams that already run Elastic, which fits a product that never opens an answer.
The ZenML guide describes the same role. It calls mature hybrid search Elasticsearch’s core strength for RAG. It adds that an organisation already running it gets vector search with no new platform to learn.
For a team starting from nothing, the counts give no reason to begin here.
Pros
- All ten models name it at least once
- Combines BM25 keyword search with vector queries and ELSER re-ranking
- Shards across nodes with built-in replication and automated recovery
Cons
- Opens none of its 23 answers, at an average position of 6.65
- Can consume more memory than purpose-built vector databases and is complex to tune
Pricing: open source. The ZenML guide lists the Elastic Cloud Hosted Standard tier at $99 a month. Best for: organisations already running Elasticsearch
8. OpenSearch
Pick OpenSearch only if the stack already runs it, because that is the conditional role the recorded answers give it.
Measured: named 18 of 50 (OpenSearch 36%), first 0 of 50 (OpenSearch 0%), average position 7.17.
OpenSearch travels with Elasticsearch. GPT-5.6 Sol names both at the same rate, and the answers give them one shared, conditional slot. GPT-5.6 Terra’s comparison answer, for example, shortlisted the dedicated engines and added OpenSearch only for teams already operating PostgreSQL, Elasticsearch or OpenSearch. No captured guide covers OpenSearch, so this entry rests on the counts and the recorded answers alone. A buyer already on OpenSearch has a reason to test it first. A buyer starting fresh gets no signal here that the models would start there, and no pricing to compare against the dedicated options.
Pros
- Four of the five GPT-5.6 Sol answers include it
- Eight of the ten models name it at least once
Cons
- Claude Fable 5 and Sonar Pro never name it
- Average position 7.17, the latest of the eight entries
- No captured guide covers it, and no public pricing is recorded
Best for: teams already operating OpenSearch, as the recorded GPT-5.6 Terra answer frames it
How the tools compare
pgvector leads the table, level with Qdrant on reach and ahead of it on first mentions. Below them the table falls in steps. Pinecone sits just behind on reach and ahead of both on first mentions. Weaviate and Milvus reach 45 answers with no first mentions. Chroma, Elasticsearch and OpenSearch form a middle band named in roughly a third to a half of answers. Eight more products appear in 10 answers or fewer.
| Vendor | Named | Share | Named first | First share | Avg position | Pricing model (as recorded) |
|---|---|---|---|---|---|---|
| pgvector | 49/50 | 98% | 16/50 | 32% | 2.82 | Free extension, pay for Postgres hosting |
| Qdrant | 49/50 | 98% | 12/50 | 24% | 2.24 | Open source, free cloud tier, usage-based |
| Pinecone | 47/50 | 94% | 19/50 | 38% | 2.04 | Managed, free tier, monthly minimums |
| Weaviate | 45/50 | 90% | 0/50 | 0% | 4.02 | Open source, serverless cloud plans |
| Milvus | 45/50 | 90% | 0/50 | 0% | 4.64 | Open source, Zilliz Cloud tiers |
| Chroma | 24/50 | 48% | 2/50 | 4% | 6.42 | Not recorded |
| Elasticsearch | 23/50 | 46% | 0/50 | 0% | 6.65 | Open source, Elastic Cloud tiers |
| OpenSearch | 18/50 | 36% | 0/50 | 0% | 7.17 | Not recorded |
| LanceDB | 10/50 | 20% | 0/50 | 0% | 7.4 | Not recorded |
| Redis | 9/50 | 18% | 0/50 | 0% | 8 | Open source, memory-based cloud plans |
| Turbopuffer | 7/50 | 14% | 0/50 | 0% | 5.43 | Managed, plan minimums |
| Vespa | 5/50 | 10% | 0/50 | 0% | 8.2 | Open source, resource-based cloud |
| Supabase | 3/50 | 6% | 0/50 | 0% | 5 | Not recorded |
| Neon | 3/50 | 6% | 0/50 | 0% | 6 | Not recorded |
| MongoDB | 3/50 | 6% | 0/50 | 0% | 8.67 | Free tier, usage-based Atlas clusters |
| FAISS | 1/50 | 2% | 0/50 | 0% | 3 | Not recorded |
Share is the share of the 50 answers that named the product. Named first counts answers where it appeared before any other tracked product. Average position is its mean place among tracked products in the answers that named it, so lower means earlier.
The pricing column comes from the ZenML guide’s plan listings, such as Pinecone’s free Starter tier, Qdrant Cloud’s free 1GB cluster and MongoDB’s usage-based Atlas clusters.
The products below the top eight each have a narrow audience in the answers, and none is named first in any answer. The guides describe five of them.
Turbopuffer is a managed vector and full-text search database built on object storage. Redis offers vector search as an extension of its in-memory engine. MongoDB builds vector search into its platform through Atlas Vector Search. Vespa works as both a search engine and a vector database. The ZenML guide lists Supabase among the Postgres-compatible services pgvector runs on.
Why does the named-first slot split three ways?
Because the models agree on the shortlist and disagree on the default. pgvector, Qdrant and Pinecone each appear in 47 to 49 answers. The first slot splits Pinecone 19, pgvector 16, Qdrant 12, and only Chroma, first in 2 of 50, breaks in.
Most answers open by declining to name one best product, then pick a default for an assumed buyer. The assumed buyer changes with the model family. GPT-5.6 Sol: “For most new RAG applications, my default recommendation is Qdrant.” Claude Opus 5: “Start with pgvector on Postgres.” Sonar Pro: “I would recommend Pinecone as the default choice if you want the safest managed, zero-ops option.” None of the five prompts states a stack, so each model fills that gap in its own way.
The defaults also track what each answer retrieved. The GPT-5.6 answers that default to Qdrant cite Qdrant’s own documentation. The Claude answers that open with pgvector cite third-party roundups, many of them on medium.com, the second most-cited host in the run at 102 citations. Claude Opus 5 flagged the weakness of that source pool in one answer: “nearly all of these are vendor blogs, SEO content marketing, or Medium listicles.”
pgvector’s reach carries one more qualifier. The panel counts an answer for pgvector when it says Postgres or PostgreSQL. GPT-5.6 Terra’s startup answer chose Qdrant Cloud and kept document and permission data in Postgres, and it counts for both. The row is an accurate count of names. It is not a count of recommendations.
The practical reading is simple. The shortlist is stable across every model tested. The default is not, and it moves with the assistant and with the pages that assistant happened to retrieve. A buyer who asks one assistant and takes its first name has sampled one family’s habit, not the category.
Where the models disagree
On the leader, nowhere. pgvector is the measured leader for all ten models and all four families. The disagreement sits below the top three.
Weaviate is the only top-five product with a clear family gap. Anthropic and Google models name it in every answer. The OpenAI family names it in Weaviate 73.3% of its 15 answers, with GPT-5.6 Terra at 3 of 5.
GPT-5.6 Sol leans toward the search-engine incumbents. It names Elasticsearch in 4 of 5 answers, as often as any model, and it is the one model that never names Chroma. Chroma’s wider split, heavy in Claude and light elsewhere, is set out in its entry.
The Gemini models carry the embedded and Postgres-hosted options. Gemini 3.6 Flash and Gemini 3.5 Flash each name LanceDB in 3 of 5 answers, and Gemini 3.6 Flash names Supabase and Neon in 2 of 5 each. No other model names Supabase or Neon except GPT-5.6 Luna, once each.
Claude Opus 5 is the model most likely to raise Turbopuffer, in 3 of 5 answers, with Claude Sonnet 5 at 2 of 5. Sonar Reasoning Pro is the only model to name Vespa and MongoDB in 2 of 5 answers each. FAISS appears once, in a GPT-5.6 Terra answer.
Sonar Reasoning Pro’s lower counts on the top three need a footnote. One of its answers stops mid-sentence before naming any product, so every vendor misses it. Its 4 of 5 on pgvector and Qdrant reflects that truncation, not a choice.
How the sample was built
10 models x 5 fixed prompts = 50 recorded answers. Each model answered each prompt once, in the 2026-09 edition, run on 2 September 2026. The panel tracked 16 vendors, and all 16 were named at least once.
The five questions, verbatim:
- What is the best vector database for a RAG application? Name specific products.
- Which vector database would you recommend to a RAG application in 2026?
- Compare the top vector database options right now.
- I’m a RAG application and I need a vector database. What should I use and why?
- Best vector database for a RAG application for a startup building AI search?
The models, by family:
- OpenAI, 15 answers: GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna
- Anthropic, 15 answers: Claude Opus 5, Claude Sonnet 5, Claude Fable 5
- Google, 10 answers: Gemini 3.6 Flash, Gemini 3.5 Flash
- Perplexity, 10 answers: Sonar Pro, Sonar Reasoning Pro
GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.6 Flash and Sonar Pro fill the panel’s ChatGPT, Claude, Gemini and Perplexity slots. The full method is on the method page, and every answer with its citations is in the vector database index record.
How this sits against the vector database guides
The guides that rank for this question sell a purchase decision or teach a setup. This page counts names in AI answers, which none of them does.
The category itself is simple to state. A vector database stores embeddings and returns the chunks a RAG app uses as context.
ZenML’s guide, dated October 1, 2025, tests ten databases and lists Pinecone first. It places pgvector eighth of the ten. It closes by pitching ZenML as the orchestration layer, arguing that integration matters more than the database chosen.
Most of the pricing on this page comes from that guide. Vendor prices change, so confirm them on the vendor’s own pricing page before budgeting.
Braintrust’s guide covers five databases: Pinecone, Weaviate, Qdrant, Chroma and Turbopuffer. It lists Milvus and pgvector as honorable mentions. It pitches Braintrust’s own evaluation product for measuring retrieval and answer quality.
The recorded answers cite braintrust.dev 39 times, so this guide is one of the pages feeding the models.
IBM’s explainer defines RAG vector databases and walks through how retrieval works. Its worked example is a Wikimedia Deutschland project built on DataStax Astra DB on IBM watsonx.data, an IBM product.
A YouTube tutorial by Francesco Ciulla builds a RAG video search on DataStax Astra DB and Langflow. Its description carries DataStax sign-up links.
A Reddit thread in the same results could not be captured.
The sharpest contrast is pgvector. Both buying guides above place it low, and the recorded answers name it more than any other product. What this page adds is that measurement, the per-model split behind it, and the finding that the first name an assistant gives depends on the assistant.
What should a buyer do with this?
Match the default to the stack you run, then test two candidates on your own data. The models agree on this much, and it is the most useful thing the counts say.
- Already on PostgreSQL: start with pgvector, and move when latency, tenant isolation, write volume or dataset size outgrow it.
- No one to run infrastructure: Pinecone.
- Open source, with filtering by tenant, permission or date: Qdrant.
- Queries full of exact names, codes or IDs: Weaviate, for hybrid search.
- Very large collections and a team to run a cluster: Milvus.
- Prototype or small app: Chroma, with a plan for where it moves later.
Gemini 3.5 Flash put the same advice in one line: “Do not over-engineer on day one.”
Then check the source behind any default an assistant gives you. The answers that name Qdrant first lean on Qdrant’s documentation, and the answers that name pgvector first lean on third-party roundups. Neither is a test on your workload.
Braintrust’s guide makes a related point: a vector database cannot confirm whether the final answer is correct or grounded in the retrieved context.
Evaluate retrieval and answers together before committing, and use the index record to read the answers and their citations for yourself.
What these counts cannot tell you
A count of names says nothing about product quality, uptime, support, pricing fairness or fit with a particular stack. Being named differs from being recommended, because an answer can list a product only to warn against it.
Each model answered each prompt once, so a single answer can swing a model’s count by one. The run is one dated snapshot, the 2026-09 edition. Names are matched by string: pgvector counts on Postgres or PostgreSQL, Milvus on Zilliz and Elasticsearch on Elastic. Answers came through the models’ APIs, which can differ from the consumer chat products. All five prompts were in English. No vendor paid to appear, be reordered or be removed.
Frequently asked questions
Which vector database is best for RAG?
No single product, on this panel’s evidence. The ten models agree on a three-way shortlist of pgvector, Qdrant and Pinecone, and they split on which to name first. The deciding question is your stack: an existing Postgres database points to pgvector, a team with no database operators points to Pinecone, and filter-heavy open-source retrieval points to Qdrant.
What db to use for RAG?
Start with the database you already run if it is PostgreSQL, and add pgvector. Braintrust’s guide calls pgvector usually enough when your vectors fit comfortably inside PostgreSQL. A dedicated vector database becomes easier to justify once search latency, tenant isolation, write volume, filtering depth or dataset size push past what Postgres handles cleanly.
Does RAG use a vector database?
Usually, but not always. IBM notes that most RAG systems rely on vector databases or vector indexing to enable semantic search. The same explainer says vector search is not strictly required, and RAG can also use keyword search, structured queries or hybrid approaches.
How to set up a vector database for RAG?
Split the source documents into chunks, embed them, and store the embeddings with metadata such as source, date, tags or permissions. At query time, embed the question, search for the nearest chunks, apply filters, and pass the top matches to the model as context. Chunk size matters: chunks that are too large lose precision, and chunks that are too small lose context. A worked example is the DataStax tutorial, which creates an Astra DB collection with an OpenAI embedding model and cosine similarity.
Can a team switch vector databases later?
Usually, without rebuilding the whole RAG app, as long as the source documents, chunks, embeddings and evaluation set live outside the database. The work is re-indexing the vectors, updating the database client and checking that retrieval quality did not regress. That makes the first choice less permanent than the guides imply.
How this list is ordered
The order is the measurement, not an assessment of the products. Answer share is the share of recorded answers that named the tool. Named first is the share where it appeared before any other tracked tool. Both are counts from one dated edition and are published in full on the category page.
A tool appears here only if it was named in the edition and its record carries a sourced claim. A product that was never named is not listed, and no position is sold.
Where to check it
- The full category record every answer, per-model split, cited sources
- The recorded answers raw output and counts
- The method how the panel runs and what is counted