Cassandra: when to use it, and why not just Postgres
Published 7 October 2026
Cassandra is a distributed NoSQL database built to store very large volumes of data across many machines, with high write throughput and continuous availability. It earns its place when the system must keep accepting reads and writes even while machines, or whole data centers, are down.
Where Cassandra is used¶
- Time-series and event data — device telemetry, app events, logs, activity feeds.
- High-volume writes — workloads that ingest huge numbers of records continuously.
- Globally distributed services — data replicated across regions so users read and write against a nearby copy.
- Predictable key-based queries — "get this user's events for the last day", where the access patterns are known up front.
Cassandra works best when each table is designed around a specific query. There are no joins and little ad hoc querying, so the question comes first and the table second:
-- Query: "events for user X in the last day, newest first"
CREATE TABLE events_by_user (
user_id uuid,
day date,
event_time timestamp,
event_type text,
payload text,
PRIMARY KEY ((user_id, day), event_time)
) WITH CLUSTERING ORDER BY (event_time DESC);
(user_id, day) is the partition key: it decides which nodes own the data. event_time is
the clustering key: it sorts rows inside the partition so the read is one sequential scan.
A different question ("all events of type X") needs a different table.
Why not Postgres?¶
Postgres is usually the better default. It has SQL, joins, constraints, transactions and flexible querying, and with careful indexing, partitioning, caching and replication it handles substantial scale.
Cassandra becomes attractive when those techniques stop being enough: very high write volume, horizontal scaling across many nodes or regions, and availability during failures. It spreads data across nodes by partition key and replicates it, so there is no single primary that every write has to go through.
The price is giving up Postgres conveniences:
| Need | Cassandra | Postgres |
|---|---|---|
| Flexible joins and ad hoc queries | Poor fit | Strong fit |
| Relational constraints and multi-row transactions | Limited | Strong |
| Scaling writes across many nodes | Built in | Needs extra architecture |
| Multi-region write availability | Designed for it | Possible, but more involved |
| Query patterns known in advance | Strong fit | Works, may be more than needed |
Rule of thumb¶
Start with Postgres unless there is a measured need for Cassandra's distributed write scale or multi-region availability. Cassandra's operational and data-modeling costs are substantial, so "big data" alone isn't a reason to pick it.