---
title: "pgvector vs Pinecone: when Postgres stops being enough"
description: "pgvector vs Pinecone: most teams reach for a dedicated vector DB too early. The thresholds - memory, QPS, filtered search - where pgvector stops being enough."
url: https://zeroic.in/blog/pgvector-vs-pinecone
---

[The Zeroic Journal](https://zeroic.in/blog)

[AI development](https://zeroic.in/blog/topic/ai-development) September 23, 2026 8 min read

# pgvector vs Pinecone: when Postgres stops being enough

The four signals that tell you pgvector has hit its limit, and the fixes to try inside Postgres before you pay for a second database.

By Prashant Abbi Partner @ Zeroic

On this page

1. [pgvector is no longer the slow option](#pgvector-is-no-longer-the-slow-option)
2. [The memory cliff: index size vs RAM](#the-memory-cliff-index-size-vs-ram)
3. [Filtered search: fixed in pgvector 0.8](#filtered-search-fixed-in-pgvector-08)
4. [When a dedicated vector database is worth it](#when-a-dedicated-vector-database-is-worth-it)
5. [How to decide](#how-to-decide)
6. [How FormulaBot serves 1.5M+ users with pgvector in the app database](#how-formulabot-serves-15m-users-with-pgvector-in-the-app-database)

On this page 6 sections

1. [pgvector is no longer the slow option](#pgvector-is-no-longer-the-slow-option)
2. [The memory cliff: index size vs RAM](#the-memory-cliff-index-size-vs-ram)
3. [Filtered search: fixed in pgvector 0.8](#filtered-search-fixed-in-pgvector-08)
4. [When a dedicated vector database is worth it](#when-a-dedicated-vector-database-is-worth-it)
5. [How to decide](#how-to-decide)
6. [How FormulaBot serves 1.5M+ users with pgvector in the app database](#how-formulabot-serves-15m-users-with-pgvector-in-the-app-database)

pgvector handles most production RAG systems, and vector count alone isn’t the signal to move to Pinecone. Postgres stops being enough when your HNSW index outgrows about half your RAM and halfvec can’t fix it, when you need sub-10ms p99 at 1,000+ queries per second, or past 50M vectors without pgvectorscale.

The fourth signal is multi-tenant: one heavy tenant slowing search for everyone else. Short of those, leaving early buys you a second service and a sync problem you didn’t need. The old advice - start on pgvector, migrate to Pinecone when you get serious - dates from when pgvector had only IVFFlat indexes and a filtered-search problem. Both are fixed.

## pgvector is no longer the slow option

The narrative that Pinecone is for production and pgvector is for small shops got locked in early and hasn’t been updated with the data.

On 50 million 768-dimension Cohere embeddings, self-hosted PostgreSQL with pgvector and pgvectorscale delivered 28x lower p95 latency and 16x higher query throughput than Pinecone’s storage-optimized pod index (s1) at 99% recall, at about 25% of the monthly cost. Against Pinecone’s performance-optimized pod index (p2) at 90% recall, the Postgres setup still came out ahead: 1.4x lower p95 latency, 1.5x higher throughput, at about 21% of the cost. Timescale, which builds pgvectorscale, ran the benchmark and published the methodology in 2024 - so read it as a vendor benchmark, but one you can reproduce.

Production teams have made the same call. Confident AI, for one, moved from Pinecone to pgvector, citing Pinecone’s metadata limits and the extra plumbing needed to join vector results back to their Postgres data.

At under 1 million vectors, the performance difference between pgvector and a dedicated database is usually smaller than the latency of the embedding API call that runs before the search. If your pipeline calls an embedding API at 50-100ms, you will not notice whether your nearest-neighbor search took 8ms or 15ms.

## The memory cliff: index size vs RAM

Vector count is a proxy. What decides whether latency stays stable is whether your HNSW index fits in RAM.

A working rule of thumb: once your HNSW index grows past roughly half of your instance’s memory, latency gets less predictable. HNSW is fast because it expects the index to be in memory. The index keeps serving queries, but when it no longer fits in cache alongside the rest of your data, more of it is fetched from disk and tail latency climbs under load. More RAM, not a different database, is the fix - until more RAM stops being affordable.

Here’s the memory math for 1,536-dimension vectors (the default output size for OpenAI’s text-embedding-3-small):

Raw vector data is `N × D × 4 bytes`, about 6 KB per vector at 1,536 dimensions. But pgvector fits only one vector of that size on each 8 KB index page, so plan on about 8 KB per vector. Jonathan Katz measured a 7,734 MB HNSW index for 1 million 1,536-dimension vectors.

- **100,000 vectors** - ~0.8 GB index. Any Postgres instance handles this without configuration changes.
- **500,000 vectors** - ~4 GB. A 16 GB instance shares it with shared_buffers and the rest of the database comfortably.
- **1,000,000 vectors** - ~8 GB. Half of a 16 GB instance, so size up to 32 GB if the same database serves your production workload. At this size, vector indexes compete with your application for memory.
- **5,000,000 vectors** - ~40 GB. Large dedicated instance. This is where you run the cost comparison against Pinecone serverless.

If you’re under 1 million vectors: start with the defaults. HNSW defaults (`m=16`, `ef_construction=64`, `hnsw.ef_search=40`) are a sensible starting point; measure recall on your own queries and raise `hnsw.ef_search` if it’s short. HNSW builds slower and uses more memory than IVFFlat, but it has no training step, so you can create the index on an empty table and insert as data arrives.

Between 1 million and 10 million vectors, tune two things:

First, raise `maintenance_work_mem` before running the index build. The Postgres default is 64 MB. pgvector’s docs note that HNSW indexes build significantly faster when the graph fits into `maintenance_work_mem`, and pgvector prints a notice when it no longer does. Set it to several GB for the build session (the docs use 8 GB as the example), and raise `max_parallel_maintenance_workers` to build in parallel.

Second, consider halfvec. pgvector 0.7.0 added native half-precision vector storage that roughly halves index size. In Jonathan Katz’s tests on 1M 1,536-dimension OpenAI embeddings, the index went from 7,734 MB to 3,867 MB with the same recall at default search settings. That’s the difference between needing a 32 GB and a 16 GB instance.

## Filtered search: fixed in pgvector 0.8

The most common complaint from serious pgvector users wasn’t raw speed. It was this: combine a vector similarity search with a `WHERE` clause - filter by user, by document type, by date range, by tenant - and pgvector’s HNSW index would exhaust its `ef_search` budget before finding enough filtered results. Filtering is applied after the index scan, so per pgvector’s docs, if a condition matches 10% of rows, the default `ef_search` of 40 returns only 4 matching rows on average. You’d request 10 similar documents and get back 4, because the index stopped scanning when it hit the limit, not when it found enough matches.

The standard workaround was to fetch far more results than needed, then filter in application code. Which removes most of the benefit of an index.

pgvector 0.8’s iterative scanning removes this workaround. When you set `hnsw.iterative_scan = 'relaxed_order'` (or `'strict_order'` for exact result ordering), the index keeps scanning until it finds enough filtered results or reaches `hnsw.max_scan_tuples` (20,000 by default), rather than stopping at the ef_search limit. For per-user, per-tenant, or per-document-type filtering - which describes most production RAG pipelines - this is the change that matters most.

If you’re on a pgvector version below 0.8 and you’ve been frustrated by filtered search quality, upgrade before you evaluate any migration. The upgrade is probably all you need.

## When a dedicated vector database is worth it

Four specific signals make a dedicated database worth the operational overhead of a separate service, a sync layer, and additional infrastructure.

**Sub-10ms p99 at 1,000+ queries per second.** It’s a narrow requirement, but some products have it. Qdrant is a purpose-built vector search engine, tuned for high-throughput approximate search in a way that PostgreSQL’s general-purpose executor isn’t. If you’re building a real-time recommendation engine or a high-throughput autocomplete feature at consumer scale, Qdrant is worth the extra service. Under 500 QPS, the difference doesn’t justify the operational cost.

**50+ million vectors without pgvectorscale.** At this scale, HNSW index rebuild times on standard pgvector become an operational event - hours, not minutes. Pinecone and Qdrant handle index management transparently. pgvectorscale closes this gap on PostgreSQL, but it’s still a Postgres extension you’re running, not a managed service. If you’re at 100M+ vectors and want zero infrastructure management, Pinecone serverless is rational.

**Per-tenant strict performance isolation at high tenant count.** Row-level security in PostgreSQL handles multi-tenant RAG cleanly at hundreds of tenants. Where it breaks down is when a single high-volume tenant can degrade search for all others by saturating the index’s cache budget. Pinecone’s serverless namespaces store each tenant separately so one tenant’s reads and writes don’t affect others, and Weaviate’s multi-tenancy puts each tenant on its own shard. RLS on a shared index can’t give you that.

**Fully managed auto-scaling with no DBA.** Pinecone serverless bills on usage (the Standard plan has a $50 monthly minimum) and scales without you touching instance sizing. If your team has no infrastructure capability and operational overhead is the binding constraint, the premium for fully managed is worth paying even at smaller vector counts.

Where a dedicated engine like Qdrant does win is the combination: tens of millions of vectors, high QPS and heavy filtering at the same time. Filter-heavy queries are what make pgvector’s p99 hard to hold at that scale, so a workload like that can justify the move even before it crosses the 1,000 QPS line.

Everything outside these signals: stay on pgvector.

## How to decide

The vector counts below assume 1,536-dimension embeddings and stand in for index size. The number you’re watching is the index-to-RAM ratio.

- **Under 1M vectors, already on Postgres.** Add pgvector, build the HNSW index with default settings, ship it. Skip the vendor bake-off; just check recall on a sample of real queries. Revisit when the index size approaches half your instance’s RAM.
- **1M-5M vectors.** Stay on pgvector. Enable halfvec to halve the index size. Tune `maintenance_work_mem` before the index build. Monitor the index-to-RAM ratio.
- **Filtered queries at any vector count on pgvector < 0.8.** Upgrade to 0.8+ and enable iterative scanning before evaluating migration. This is the most likely fix.
- **5M-50M vectors.** Evaluate pgvectorscale before evaluating Pinecone. It’s a Postgres extension - no new infrastructure, no sync layer - and published benchmarks at 50M vectors show it beating Pinecone’s storage tier on cost, latency, and throughput simultaneously.
- **50M+ vectors OR sub-10ms p99 at 1,000+ QPS.** Now run the benchmark seriously. Qdrant for open-source speed with strong filtering; Weaviate for hybrid search and multi-tenancy; Pinecone for fully managed ops.
- **Already paying for Pinecone at under 5M vectors.** There’s a cost case for moving to pgvector. Don’t migrate for its own sake - but if you’re re-architecting or onboarding new workloads, default to Postgres first.
- **Multi-tenant, hundreds of tenants, moderate vector counts.** Postgres row-level security handles this cleanly. Switch to per-namespace isolation only when a noisy tenant is measurably degrading others, which you’ll see in p99 latency by tenant segment.

Not sure which bucket you’re in? An [AI Audit](https://zeroic.in/services/ai-audit) maps your retrieval setup, costs and scaling limits in one to two weeks - or [book a call](https://zeroic.in/book) and talk it through with the engineer who runs ours.

## How FormulaBot serves 1.5M+ users with pgvector in the app database

At [FormulaBot / Better Analyst](https://zeroic.in/formulabot), the RAG pipeline lives on pgvector in the same PostgreSQL instance as the application database. The semantic search layer handles three jobs: similar-question lookup to retrieve answers to semantically equivalent past queries, schema context retrieval to pull relevant table descriptions into the LLM prompt, and query-result caching keyed by embedding similarity. The product serves 1.5M+ users across 40+ data connectors, and the HNSW indexes still share one Postgres instance with the application data.

You can read more about the kinds of [AI systems we build on our work page](https://zeroic.in/work).

## Frequently asked questions

**Is pgvector slower than Pinecone?**

Usually not in a way you'd notice. Under 1 million vectors, the search is a small part of request latency next to the embedding API call before it. In Timescale's 2024 benchmark on 50 million 768-dimension vectors, pgvector with pgvectorscale showed 28x lower p95 latency than Pinecone's storage-optimized s1 pod index at 99% recall, for about a quarter of the cost. Timescale builds pgvectorscale, so treat it as a vendor benchmark.

**How much RAM does a pgvector HNSW index need?**

Plan on about 8 KB per vector for 1,536-dimension embeddings, because pgvector fits only one vector of that size on each 8 KB index page. Jonathan Katz measured a 7,734 MB HNSW index for 1 million of them. A working rule of thumb: once the index grows past roughly half your instance's RAM, it competes with the rest of the database for cache and tail latency gets less predictable.

**Does pgvector support filtered vector search?**

Yes, properly since pgvector 0.8. Before it, filtering ran after the index scan, so HNSW returned at most ef_search candidates (40 by default) before the WHERE clause, even if too few of them matched. Version 0.8 added iterative index scans (the hnsw.iterative_scan setting), which keep scanning until enough rows match or hnsw.max_scan_tuples (20,000 by default) is reached. If you're on an older version, upgrade before you consider migrating.

**At what vector count should I switch to Pinecone or Qdrant?**

Vector count alone isn't the right signal. The signals we'd watch are an HNSW index that has grown past roughly half your RAM and can't be fixed with a bigger instance or halfvec, QPS requirements above 1,000 with sub-10ms p99 latency, 50+ million vectors without pgvectorscale, or one tenant's load slowing search for everyone else. Short of these, pgvector almost always wins on cost-performance.

**Can pgvector handle multi-tenant RAG?**

Yes, with row-level security in Postgres. Filter by tenant_id on every query and RLS policies ensure each tenant only searches their own vectors. This pattern holds cleanly at hundreds of tenants. The limit is when a noisy high-volume tenant degrades search for others, because every tenant shares one index. If you need strict per-tenant isolation at thousands of tenants, Pinecone's serverless namespaces (each stored separately) or Weaviate's multi-tenancy (each tenant on its own shard) give you that.

**What is halfvec in pgvector and should I use it?**

halfvec, added in pgvector 0.7.0, stores vectors as 2-byte half-precision floats instead of 4-byte floats, roughly halving the index size. In Jonathan Katz's tests on 1 million 1,536-dimension OpenAI embeddings, the HNSW index went from 7,734 MB to 3,867 MB with the same 96.8% recall at the default ef_search. If you're approaching your RAM ceiling, use it after checking recall on your own data. It roughly doubles your capacity on the same instance.

## References

1. [pgvector - Open-source vector similarity search for Postgres](https://github.com/pgvector/pgvector)
2. [pgvector changelog](https://github.com/pgvector/pgvector/blob/master/CHANGELOG.md)
3. [pgvector 0.8.0 - iterative index scans - Nile](https://www.thenile.dev/blog/pgvector-080)
4. [HNSW Indexes with Postgres and pgvector - Crunchy Data](https://www.crunchydata.com/blog/hnsw-indexes-with-postgres-and-pgvector)
5. [Scalar and binary quantization for pgvector vector search and storage - Jonathan Katz](https://jkatz05.com/post/postgres/pgvector-scalar-binary-quantization/)
6. [Pgvector Is Now Faster than Pinecone at 75% Less Cost - Timescale](https://www.tigerdata.com/blog/pgvector-is-now-as-fast-as-pinecone-at-75-less-cost)
7. [Why we replaced Pinecone with PGVector - Confident AI](https://www.confident-ai.com/blog/why-we-replaced-pinecone-with-pgvector)
8. [Implement multitenancy - Pinecone docs](https://docs.pinecone.io/guides/index-data/implement-multitenancy)
9. [Multi-tenancy - Weaviate docs](https://docs.weaviate.io/weaviate/manage-collections/multi-tenancy)
10. [Pinecone pricing](https://www.pinecone.io/pricing/)
11. [VectorChord vs pgvector vs pgvectorscale - memory and disk comparison](https://blog.vectorchord.ai/vector-search-over-postgresql-a-comparative-analysis-of-memory-and-disk-solutions)
12. [The Case Against pgvector - Alex Jacobs](https://alex-jacobs.com/posts/the-case-against-pgvector/)

#pgvector#pinecone#rag#vector-search#postgres#embeddings

More from the journal

[ AI development

## [GPT vs Claude vs Llama in production: choose by task type](https://zeroic.in/blog/gpt-vs-claude-vs-llama-production)

Four kinds of LLM work, the model tier that wins each one, and the volume at which running open weights yourself starts to make sense.

Prashant Abbi Sep 29, 2026 · 8 min read

[ AI development

## [AI voice agent architecture: the four parts that decide if it ships](https://zeroic.in/blog/ai-voice-agent-architecture)

Where each layer of a production voice agent breaks on live phone calls, and the order we'd build them in, from voice agents we run in production.

Prashant Abbi Sep 27, 2026 · 10 min read

[ AI development

## [English-to-SQL at a million users: what breaks, in order](https://zeroic.in/blog/english-to-sql-at-1m-users)

The five things that broke as FormulaBot's English-to-SQL pipeline grew to 1.5M+ users, and the order we'd build the fixes in.

Prashant Abbi Sep 19, 2026 · 9 min read

## Hear from Prashant within 24 hours.

**Prashant Abbi** Partner @ Zeroic

Not sure if your vector setup will scale? Our AI Audit evaluates your retrieval architecture, tells you which thresholds you're approaching, and surfaces the specific fixes before you hit a production incident. Fixed scope, one to two weeks.

[See AI audits](https://zeroic.in/services/ai-audit), or [book a 30-minute call](https://zeroic.in/book)

Pages [Services](https://zeroic.in/services) [Work](https://zeroic.in/work) [Blog](https://zeroic.in/blog) [About](https://zeroic.in/about) [Book a call](https://zeroic.in/book)

Services [Audits](https://zeroic.in/services/ai-audit) [MVP Sprint](https://zeroic.in/services/mvp-development) [Retainer](https://zeroic.in/services/senior-developers) [Code Rescue](https://zeroic.in/services/code-rescue) [Bubble to code](https://zeroic.in/bubble-to-code) [AI development](https://zeroic.in/ai-development-company)

Contact [Email](mailto:zeroic.in@gmail.com) [WhatsApp](https://wa.me/919873927225) [LinkedIn](https://www.linkedin.com/company/zeroic/) [X (Twitter)](https://x.com/PrashantAbbi) [Guide for AI](https://zeroic.in/llms.txt)

© 2026 Zeroic [Privacy](https://zeroic.in/privacy) [Terms](https://zeroic.in/terms)

Built in India · Shipping globally

SINCE
