MemetikEdition 2026-09

Lists / AI infra

Best vector databases for AI apps (2026): What ChatGPT, Claude & Gemini Recommend

pgvector, Qdrant and Pinecone lead 50 recorded AI answers on vector databases for RAG. A ranked shortlist with who should pick each, the per-model split and pricing.

pgvector, named in 49 of 50 recorded AI answers and first in 16 of 50, is the vector database the models name most for RAG. Qdrant matches its reach. Pinecone opens more answers, first in 19 of 50. Those three are the shortlist, and the right pick depends on the stack you already run.

This page counts which products 10 AI models name when asked about vector databases for RAG, in the 2026-09 edition. It does not test the products.

TL;DR

1. pgvector

Pick pgvector if your application already runs on PostgreSQL and you want embeddings stored next to the data they describe.

Measured: named 49 of 50 (pgvector 98%), first 16 of 50 (pgvector 32%), average position 2.82.

pgvector is the measured leader for all ten models. Its only miss is a Sonar Reasoning Pro response that stops mid-sentence before it names any product. The answers that recommend it usually attach a condition, a team that already runs Postgres, and the section on the first slot below explains why its lead is softer than the row suggests.

The ZenML guide explains the appeal. If an application already uses Postgres, pgvector keeps data and vectors in one database, with full SQL querying power on the same data that backs the app. The same guide warns that Postgres isn’t designed for massive vector workloads.

That’s the trade a buyer accepts for running one database instead of two.

Pros

Cons

Pricing: free open-source extension. Costs are for PostgreSQL hosting. Best for: teams already on PostgreSQL that want to add RAG without a new database

2. Qdrant

Pick Qdrant if retrieval is the core of the product and results must be narrowed by tenant, permission or date before the final chunks are returned.

Measured: named 49 of 50 (Qdrant 98%), first 12 of 50 (Qdrant 24%), average position 2.24.

Qdrant ties pgvector on reach and misses the same truncated answer. The difference is who makes it the default. The GPT-5.6 answers most often do, and when they do they cite Qdrant’s own documentation. The host qdrant.tech is the most-cited vendor-owned host in the run, at 33 citations. Its average position is earlier than pgvector’s, so the answers tend to reach it sooner even when they open with something else.

Pros

Cons

Pricing: open source. Qdrant Cloud offers a free 1GB cluster, then usage-based managed clusters. Best for: filter-heavy retrieval

3. Pinecone

Pick Pinecone if nobody on the team should be running database infrastructure.

Measured: named 47 of 50 (Pinecone 94%), first 19 of 50 (Pinecone 38%), average position 2.04.

Pinecone trails the top two on reach and leads them on position. The gap on reach comes mostly from the OpenAI family (Pinecone 86.7% of its 15 answers). The answers that lead with it treat it as the managed pick.

Braintrust’s guide describes the same role. Pinecone runs serverless indexes as its default architecture, and its bring-your-own-cloud option can run the data plane in the buyer’s own cloud account. The guide frames managed hosting as scaling without owning upgrades, backups, monitoring and capacity planning.

For a team without database operators, that’s the whole case.

Pros

Cons

Pricing: free Starter tier. Standard carries a $50 monthly minimum and Enterprise a $500 monthly minimum. Best for: managed RAG with low infrastructure work

4. Weaviate

Pick Weaviate if hybrid keyword-and-vector search is central to the application and the team wants an open-source base.

Measured: named 45 of 50 (Weaviate 90%), first 0 of 50 (Weaviate 0%), average position 4.02.

Weaviate appears in 45 answers and opens none of them. Its 90-point gap between named and first is the widest in the table, shared only with Milvus. The models treat it as a specialist. They reach for it once the answer turns to hybrid search or multi-tenant setups, not as an opening default.

Braintrust’s guide gives the reason. Weaviate combines BM25 keyword search with vector search and lets teams tune how the two result sets are fused. Hybrid search helps with exact strings such as product names, error codes and ticket IDs.

A buyer whose queries depend on strings like those has the clearest reason to test Weaviate.

Pros

Cons

Pricing: open source. Weaviate Cloud serverless plans start at $25 a month. Best for: open-source RAG with hybrid search

5. Milvus

Pick Milvus if the dataset is very large and the team is prepared to run distributed vector search.

Measured: named 45 of 50 (Milvus 90%), first 0 of 50 (Milvus 0%), average position 4.64.

Milvus matches Weaviate on reach and on the empty first column, with a later average position. The answers file it under scale, as the option for billion-vector or distributed deployments once the default no longer fits. Its soft spot is Perplexity. Sonar Reasoning Pro names it in 3 of 5 answers.

The ZenML guide gives the reason for the scale role. Milvus separates compute and storage in a distributed architecture. It is a graduate project of the LF AI & Data Foundation. IBM’s RAG explainer names it beside Pinecone as a platform teams build with.

Pros

Cons

Pricing: open source. Zilliz Cloud has a free tier, Dedicated plans from $99 a month and Serverless at $0.30 per GB per month. Best for: very large collections and distributed deployments

6. Chroma

Pick Chroma to get a prototype or small RAG app retrieving quickly, before committing to production infrastructure.

Measured: named 24 of 50 (Chroma 48%), first 2 of 50 (Chroma 4%), average position 6.42.

Chroma carries the sharpest model split in the category. Claude Fable 5 names it in 4 of 5 answers. GPT-5.6 Terra, Gemini 3.6 Flash, Sonar Pro and Sonar Reasoning Pro name it once each. A buyer’s chance of hearing about Chroma depends mostly on which assistant they ask. The Claude answers place it in the early-stage slot, paired with pgvector as the move-fast option for small projects. Its two first mentions are the only ones outside the top three.

Braintrust’s guide calls Chroma the lightweight option for local development, prototypes and smaller applications that need vector search without a separate database service.

Pros

Cons

Pricing: no public pricing is recorded. Best for: prototypes and smaller apps

7. Elasticsearch

Pick Elasticsearch if the organisation already runs it and wants vector search without a new platform.

Measured: named 23 of 50 (Elasticsearch 46%), first 0 of 50 (Elasticsearch 0%), average position 6.65.

Elasticsearch appears in almost half the answers but thinly for most models. GPT-5.6 Sol and Claude Opus 5 name it in 4 of 5. Claude Sonnet 5, Claude Fable 5 and Sonar Pro name it once each. Where the answers explain it, they present it as the option for teams that already run Elastic, which fits a product that never opens an answer.

The ZenML guide describes the same role. It calls mature hybrid search Elasticsearch’s core strength for RAG. It adds that an organisation already running it gets vector search with no new platform to learn.

For a team starting from nothing, the counts give no reason to begin here.

Pros

Cons

Pricing: open source. The ZenML guide lists the Elastic Cloud Hosted Standard tier at $99 a month. Best for: organisations already running Elasticsearch

8. OpenSearch

Pick OpenSearch only if the stack already runs it, because that is the conditional role the recorded answers give it.

Measured: named 18 of 50 (OpenSearch 36%), first 0 of 50 (OpenSearch 0%), average position 7.17.

OpenSearch travels with Elasticsearch. GPT-5.6 Sol names both at the same rate, and the answers give them one shared, conditional slot. GPT-5.6 Terra’s comparison answer, for example, shortlisted the dedicated engines and added OpenSearch only for teams already operating PostgreSQL, Elasticsearch or OpenSearch. No captured guide covers OpenSearch, so this entry rests on the counts and the recorded answers alone. A buyer already on OpenSearch has a reason to test it first. A buyer starting fresh gets no signal here that the models would start there, and no pricing to compare against the dedicated options.

Pros

Cons

Best for: teams already operating OpenSearch, as the recorded GPT-5.6 Terra answer frames it

How the tools compare

pgvector leads the table, level with Qdrant on reach and ahead of it on first mentions. Below them the table falls in steps. Pinecone sits just behind on reach and ahead of both on first mentions. Weaviate and Milvus reach 45 answers with no first mentions. Chroma, Elasticsearch and OpenSearch form a middle band named in roughly a third to a half of answers. Eight more products appear in 10 answers or fewer.

Vendor Named Share Named first First share Avg position Pricing model (as recorded)
pgvector 49/50 98% 16/50 32% 2.82 Free extension, pay for Postgres hosting
Qdrant 49/50 98% 12/50 24% 2.24 Open source, free cloud tier, usage-based
Pinecone 47/50 94% 19/50 38% 2.04 Managed, free tier, monthly minimums
Weaviate 45/50 90% 0/50 0% 4.02 Open source, serverless cloud plans
Milvus 45/50 90% 0/50 0% 4.64 Open source, Zilliz Cloud tiers
Chroma 24/50 48% 2/50 4% 6.42 Not recorded
Elasticsearch 23/50 46% 0/50 0% 6.65 Open source, Elastic Cloud tiers
OpenSearch 18/50 36% 0/50 0% 7.17 Not recorded
LanceDB 10/50 20% 0/50 0% 7.4 Not recorded
Redis 9/50 18% 0/50 0% 8 Open source, memory-based cloud plans
Turbopuffer 7/50 14% 0/50 0% 5.43 Managed, plan minimums
Vespa 5/50 10% 0/50 0% 8.2 Open source, resource-based cloud
Supabase 3/50 6% 0/50 0% 5 Not recorded
Neon 3/50 6% 0/50 0% 6 Not recorded
MongoDB 3/50 6% 0/50 0% 8.67 Free tier, usage-based Atlas clusters
FAISS 1/50 2% 0/50 0% 3 Not recorded

Share is the share of the 50 answers that named the product. Named first counts answers where it appeared before any other tracked product. Average position is its mean place among tracked products in the answers that named it, so lower means earlier.

The pricing column comes from the ZenML guide’s plan listings, such as Pinecone’s free Starter tier, Qdrant Cloud’s free 1GB cluster and MongoDB’s usage-based Atlas clusters.

The products below the top eight each have a narrow audience in the answers, and none is named first in any answer. The guides describe five of them.

Turbopuffer is a managed vector and full-text search database built on object storage. Redis offers vector search as an extension of its in-memory engine. MongoDB builds vector search into its platform through Atlas Vector Search. Vespa works as both a search engine and a vector database. The ZenML guide lists Supabase among the Postgres-compatible services pgvector runs on.

Why does the named-first slot split three ways?

Because the models agree on the shortlist and disagree on the default. pgvector, Qdrant and Pinecone each appear in 47 to 49 answers. The first slot splits Pinecone 19, pgvector 16, Qdrant 12, and only Chroma, first in 2 of 50, breaks in.

Most answers open by declining to name one best product, then pick a default for an assumed buyer. The assumed buyer changes with the model family. GPT-5.6 Sol: “For most new RAG applications, my default recommendation is Qdrant.” Claude Opus 5: “Start with pgvector on Postgres.” Sonar Pro: “I would recommend Pinecone as the default choice if you want the safest managed, zero-ops option.” None of the five prompts states a stack, so each model fills that gap in its own way.

The defaults also track what each answer retrieved. The GPT-5.6 answers that default to Qdrant cite Qdrant’s own documentation. The Claude answers that open with pgvector cite third-party roundups, many of them on medium.com, the second most-cited host in the run at 102 citations. Claude Opus 5 flagged the weakness of that source pool in one answer: “nearly all of these are vendor blogs, SEO content marketing, or Medium listicles.”

pgvector’s reach carries one more qualifier. The panel counts an answer for pgvector when it says Postgres or PostgreSQL. GPT-5.6 Terra’s startup answer chose Qdrant Cloud and kept document and permission data in Postgres, and it counts for both. The row is an accurate count of names. It is not a count of recommendations.

The practical reading is simple. The shortlist is stable across every model tested. The default is not, and it moves with the assistant and with the pages that assistant happened to retrieve. A buyer who asks one assistant and takes its first name has sampled one family’s habit, not the category.

Where the models disagree

On the leader, nowhere. pgvector is the measured leader for all ten models and all four families. The disagreement sits below the top three.

Weaviate is the only top-five product with a clear family gap. Anthropic and Google models name it in every answer. The OpenAI family names it in Weaviate 73.3% of its 15 answers, with GPT-5.6 Terra at 3 of 5.

GPT-5.6 Sol leans toward the search-engine incumbents. It names Elasticsearch in 4 of 5 answers, as often as any model, and it is the one model that never names Chroma. Chroma’s wider split, heavy in Claude and light elsewhere, is set out in its entry.

The Gemini models carry the embedded and Postgres-hosted options. Gemini 3.6 Flash and Gemini 3.5 Flash each name LanceDB in 3 of 5 answers, and Gemini 3.6 Flash names Supabase and Neon in 2 of 5 each. No other model names Supabase or Neon except GPT-5.6 Luna, once each.

Claude Opus 5 is the model most likely to raise Turbopuffer, in 3 of 5 answers, with Claude Sonnet 5 at 2 of 5. Sonar Reasoning Pro is the only model to name Vespa and MongoDB in 2 of 5 answers each. FAISS appears once, in a GPT-5.6 Terra answer.

Sonar Reasoning Pro’s lower counts on the top three need a footnote. One of its answers stops mid-sentence before naming any product, so every vendor misses it. Its 4 of 5 on pgvector and Qdrant reflects that truncation, not a choice.

How the sample was built

10 models x 5 fixed prompts = 50 recorded answers. Each model answered each prompt once, in the 2026-09 edition, run on 2 September 2026. The panel tracked 16 vendors, and all 16 were named at least once.

The five questions, verbatim:

  1. What is the best vector database for a RAG application? Name specific products.
  2. Which vector database would you recommend to a RAG application in 2026?
  3. Compare the top vector database options right now.
  4. I’m a RAG application and I need a vector database. What should I use and why?
  5. Best vector database for a RAG application for a startup building AI search?

The models, by family:

GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.6 Flash and Sonar Pro fill the panel’s ChatGPT, Claude, Gemini and Perplexity slots. The full method is on the method page, and every answer with its citations is in the vector database index record.

How this sits against the vector database guides

The guides that rank for this question sell a purchase decision or teach a setup. This page counts names in AI answers, which none of them does.

The category itself is simple to state. A vector database stores embeddings and returns the chunks a RAG app uses as context.

ZenML’s guide, dated October 1, 2025, tests ten databases and lists Pinecone first. It places pgvector eighth of the ten. It closes by pitching ZenML as the orchestration layer, arguing that integration matters more than the database chosen.

Most of the pricing on this page comes from that guide. Vendor prices change, so confirm them on the vendor’s own pricing page before budgeting.

Braintrust’s guide covers five databases: Pinecone, Weaviate, Qdrant, Chroma and Turbopuffer. It lists Milvus and pgvector as honorable mentions. It pitches Braintrust’s own evaluation product for measuring retrieval and answer quality.

The recorded answers cite braintrust.dev 39 times, so this guide is one of the pages feeding the models.

IBM’s explainer defines RAG vector databases and walks through how retrieval works. Its worked example is a Wikimedia Deutschland project built on DataStax Astra DB on IBM watsonx.data, an IBM product.

A YouTube tutorial by Francesco Ciulla builds a RAG video search on DataStax Astra DB and Langflow. Its description carries DataStax sign-up links.

A Reddit thread in the same results could not be captured.

The sharpest contrast is pgvector. Both buying guides above place it low, and the recorded answers name it more than any other product. What this page adds is that measurement, the per-model split behind it, and the finding that the first name an assistant gives depends on the assistant.

What should a buyer do with this?

Match the default to the stack you run, then test two candidates on your own data. The models agree on this much, and it is the most useful thing the counts say.

  1. Already on PostgreSQL: start with pgvector, and move when latency, tenant isolation, write volume or dataset size outgrow it.
  2. No one to run infrastructure: Pinecone.
  3. Open source, with filtering by tenant, permission or date: Qdrant.
  4. Queries full of exact names, codes or IDs: Weaviate, for hybrid search.
  5. Very large collections and a team to run a cluster: Milvus.
  6. Prototype or small app: Chroma, with a plan for where it moves later.

Gemini 3.5 Flash put the same advice in one line: “Do not over-engineer on day one.”

Then check the source behind any default an assistant gives you. The answers that name Qdrant first lean on Qdrant’s documentation, and the answers that name pgvector first lean on third-party roundups. Neither is a test on your workload.

Braintrust’s guide makes a related point: a vector database cannot confirm whether the final answer is correct or grounded in the retrieved context.

Evaluate retrieval and answers together before committing, and use the index record to read the answers and their citations for yourself.

What these counts cannot tell you

A count of names says nothing about product quality, uptime, support, pricing fairness or fit with a particular stack. Being named differs from being recommended, because an answer can list a product only to warn against it.

Each model answered each prompt once, so a single answer can swing a model’s count by one. The run is one dated snapshot, the 2026-09 edition. Names are matched by string: pgvector counts on Postgres or PostgreSQL, Milvus on Zilliz and Elasticsearch on Elastic. Answers came through the models’ APIs, which can differ from the consumer chat products. All five prompts were in English. No vendor paid to appear, be reordered or be removed.

Frequently asked questions

Which vector database is best for RAG?

No single product, on this panel’s evidence. The ten models agree on a three-way shortlist of pgvector, Qdrant and Pinecone, and they split on which to name first. The deciding question is your stack: an existing Postgres database points to pgvector, a team with no database operators points to Pinecone, and filter-heavy open-source retrieval points to Qdrant.

What db to use for RAG?

Start with the database you already run if it is PostgreSQL, and add pgvector. Braintrust’s guide calls pgvector usually enough when your vectors fit comfortably inside PostgreSQL. A dedicated vector database becomes easier to justify once search latency, tenant isolation, write volume, filtering depth or dataset size push past what Postgres handles cleanly.

Does RAG use a vector database?

Usually, but not always. IBM notes that most RAG systems rely on vector databases or vector indexing to enable semantic search. The same explainer says vector search is not strictly required, and RAG can also use keyword search, structured queries or hybrid approaches.

How to set up a vector database for RAG?

Split the source documents into chunks, embed them, and store the embeddings with metadata such as source, date, tags or permissions. At query time, embed the question, search for the nearest chunks, apply filters, and pass the top matches to the model as context. Chunk size matters: chunks that are too large lose precision, and chunks that are too small lose context. A worked example is the DataStax tutorial, which creates an Astra DB collection with an OpenAI embedding model and cosine similarity.

Can a team switch vector databases later?

Usually, without rebuilding the whole RAG app, as long as the source documents, chunks, embeddings and evaluation set live outside the database. The work is re-indexing the vectors, updating the database client and checking that retrieval quality did not regress. That makes the first choice less permanent than the guides imply.

How this list is ordered

The order is the measurement, not an assessment of the products. Answer share is the share of recorded answers that named the tool. Named first is the share where it appeared before any other tracked tool. Both are counts from one dated edition and are published in full on the category page.

A tool appears here only if it was named in the edition and its record carries a sourced claim. A product that was never named is not listed, and no position is sold.

Where to check it