Design Slack
Published 7 October 2026
Design Slack.

Message IDs: Snowflake is a library, not a service¶
The diagram draws an "ID Generator (Snowflake)" box, but there is no ID service to call. Snowflake is a scheme, and the generator is a library linked into the Message Service. Each instance builds a 64-bit ID locally:
| Bits | Field | Purpose |
|---|---|---|
| 1 | sign | always 0 |
| 41 | timestamp (ms since a custom epoch) | makes IDs time-ordered |
| 10 | machine / worker ID | keeps instances from colliding |
| 12 | sequence | up to 4096 IDs per ms per instance |
So getting a msg_id is a function call with no network hop. The IDs are unique and
monotonically increasing per generator. Across instances they are roughly time-ordered,
within clock skew, which is good enough to sort a conversation and to ask "everything
after msg_id X".
Write path (A sends a message to B)¶
- A holds a WebSocket to Chat Server 1, through the L4 load balancer. The JWT from the
Auth service was validated when the socket opened, and Chat Server 1 registered
A -> chat-server-1in the Redis session store. - A sends the message over the socket with a
request_id(a UUID generated on the client) as an idempotency key. If the ack is lost and A retries with the samerequest_id, the server returns the originalmsg_idinstead of storing a duplicate. - Chat Server 1 forwards it to the Message Service, which assigns a Snowflake
msg_id. - The Message Service persists it to Cassandra, partitioned by
conv_idand sorted bymsg_id. Once the write succeeds, A gets a "sent" ack. - The Message Service looks up B in Redis:
- B online: deliver through B's chat server (Chat Server 2), and B acks back ("delivered").
- B offline: publish to the offline queue (Kafka). The Push Notification Service consumes it and sends a notification through APNs / FCM.
Media doesn't go through the socket. A uploads the file over HTTPS (API Gateway -> Media Service -> S3) and the message carries only the media URL.
Read path¶
- Live messages arrive pushed over the WebSocket; B doesn't poll.
-
Opening a conversation reads history from Cassandra: one partition (
conv_id), newest first, paginated bymsg_id: -
Media is downloaded from the CDN using the URL in the message, not from the chat servers.
Offline sync¶
The push notification only tells B there is something new. The messages themselves come back when B reconnects:
- B opens a new WebSocket (possibly to a different chat server), the JWT is validated, and the session store is updated to point at the new server.
- B sends the last
msg_idit received. - The server replays everything newer from the Message DB. Since
msg_idis time-ordered, this is a range scan:msg_id > last_receivedper conversation. - B acks what it received, and the server advances B's sync cursor, so a reconnect halfway through just resumes from the last ack.
Delivery is at least once: a replay can resend a message B already has, and B drops it by
msg_id.
Heartbeats over the socket keep presence (online / last_seen) fresh in Redis. When they
stop, the session entry expires and B counts as offline, so new messages take the push path.
Pending questions¶
- Why do we need an L4 load balancer? Can't we use L7, and what's the advantage of each?
- Can WebSocket connections be relayed through L4 load balancers?
- How does offline sync actually work? How does the server know which conversations to sync, and what if the user has thousands of chats?