Skip to content

Postgres vs MySQL: what is a connection?

Published 6 October 2026

Uber ran close to 10,000 connections on some MySQL instances. On Postgres, they struggled past a few hundred.

Same SQL. A different answer to one question: what is a connection?

Postgres process-per-connection versus MySQL thread-per-connection On the Postgres side, three clients connect to the postmaster, which forks one backend process per client. Each backend has its own private memory and reaches shared_buffers only through an explicit shared memory segment. On the MySQL side, the same three clients each get a thread inside a single mysqld process, and all threads share the process memory, including the buffer pool. Postgres · process per connection MySQL · thread per connection client 1 client 2 client 3 postmaster accepts, then fork()s backend 1 own process backend 2 own process backend 3 own process shared_buffers explicit shared memory segment client 1 client 2 client 3 mysqld · one OS process thread 1 same address space thread 2 same address space thread 3 same address space buffer pool, caches, heap all of it shared by every thread 1 · Postgres: every client connects to the postmaster 2 · fork(): one whole OS process per connection 3 · Each backend's private memory is walled off from the others 4 · MySQL: same three clients, three threads in one process 5 · Threads share everything: cheap, but a stray pointer hits all

Click the diagram to pause or play.

Postgres: one OS process per connection

Every client gets its own backend, forked by the postmaster on connect.

  • Isolation. A stray pointer in one backend can't touch another backend's private memory. Backends only meet through an explicit shared memory segment (shared_buffers, lock tables), so one bad connection crashing is a contained event.
  • Cost. A fork isn't free, and every connection, idle or not, is a process holding its own memory: catalog caches, work_mem for sorts and hashes, its stack.

MySQL: one process, one thread per connection

mysqld is a single process. Each connection is a thread inside it.

  • Cost. Threads are cheap to create, and they share memory by default — the buffer pool, caches, everything in the heap.
  • Isolation. That cuts both ways. A stray pointer can corrupt memory another connection is using, because there's no wall between them.

Neither is wrong

Postgres MySQL
Unit per connection OS process thread
Memory between connections private, plus explicit shared segment shared by default
Cost of an idle connection a whole process a thread
Blast radius of a memory bug one backend the whole server
What it optimises for isolation cost

Postgres chose isolation. MySQL chose cost.

Why pgbouncer sits in front of almost every Postgres

This is why a pooler like pgbouncer sits in front of almost every Postgres in production: thousands of clients, multiplexed onto a small set of real backends. In transaction mode a client only holds a backend for the length of one transaction, so a few dozen backends can serve thousands of mostly idle clients.

Why it matters to you

The pool is per pod, so the database sees the product, not the setting:

Setup Client pool MySQL sees Postgres sees
10 pods 20 each 200 threads 200 processes
autoscale to 50 pods 20 each 1,000 threads 1,000 processes, forked one by one

Go's database/sql won't stop you. SetMaxOpenConns has no cap by default:

db, _ := sql.Open("pgx", dsn)
db.SetMaxOpenConns(20)            // default 0 = unlimited
db.SetMaxIdleConns(20)
db.SetConnMaxIdleTime(5 * time.Minute)

Size the per-pod pool against max pods × pool size, not against one pod.

The transport layer is the first box in every database diagram. Most people never look at it. Until it's the bottleneck.

Critique: what this framing leaves out

The process-vs-thread story is a clean hook, but it's not the whole story, and a fair amount of pushback on it is worth keeping.

  • The Uber post is old. It's from 2016, and it's been rebutted many times since. Its bigger complaints were write amplification from Postgres's MVCC and index layout, and replication behaviour — not connection count. Taking "Uber left because of 10k connections" at face value overstates it; a well-run Postgres with a pooler handles very high TPS and very large client counts without trouble.
  • 10,000 connections isn't 10,000 queries running. A 64-core box can only execute about 64 things at once. Past that, more connections — threads or processes — just means context switches, lock contention and cold CPU caches. MySQL behind a pool capped near the core count will usually beat MySQL with 10k live threads. The number to reason about is concurrent execution, not pool size.
  • Postgres expects pooling by design. The intended OLTP shape is one connection per concurrent transaction. Treating that as a defect misreads the model: the pooler is part of the architecture, not a band-aid.
  • fork() on Linux is cheap. Process creation is genuinely expensive on Windows, which lacks an optimised fork(). On Linux a process is roughly a thread with its own address space, and pages are copy-on-write, so spawning is cheap. Teardown costs more than a thread exit, but much of it is amortised. The real per-connection cost is the memory a long-lived backend accumulates, not the fork itself.
  • Hold time beats pool size. On Kubernetes the multiplication is real, but the usual root cause of exhausted pools is how long each connection is held. A Kafka publish or HTTP call inside a transaction keeps a connection busy for the whole round trip and will choke any pool. Short transactions, small pools and a pooler in front tend to do more than tuning. Postgres-compatible distributed databases that keep process-per-connection don't make this go away, and can add a network hop per statement.
  • In Go, the driver matters too. For Postgres, pgx with pgxpool instead of database/sql gets you automatic prepared statements, the binary wire protocol, and native types like slices as parameters. That's often a bigger win than anything in the connection model.
  • Isolation has real value. For many workloads a memory bug taking down one backend instead of the whole server isn't a nice-to-have, and with a pooler the scalability cost of that isolation is small.
  • The broader comparison is elsewhere. Postgres's harder scaling problems tend to be write amplification, replication modes and disk management. MySQL leans towards ease of administration; Postgres towards features and SQL expressiveness.

So the honest summary: the connection model is a real difference and the Kubernetes math is a real trap, but it's a design input to plan around with a pooler, not a reason one database loses.

References