Skip to content

Design Slack

Published 7 October 2026

Design Slack.

1:1 chat HLD: auth and media over HTTPS, WebSocket chat servers, Redis session store, Cassandra message DB, Kafka offline queue and push notifications

Message IDs: Snowflake is a library, not a service

The diagram draws an "ID Generator (Snowflake)" box, but there is no ID service to call. Snowflake is a scheme, and the generator is a library linked into the Message Service. Each instance builds a 64-bit ID locally:

Bits Field Purpose
1 sign always 0
41 timestamp (ms since a custom epoch) makes IDs time-ordered
10 machine / worker ID keeps instances from colliding
12 sequence up to 4096 IDs per ms per instance

So getting a msg_id is a function call with no network hop. The IDs are unique and monotonically increasing per generator. Across instances they are roughly time-ordered, within clock skew, which is good enough to sort a conversation and to ask "everything after msg_id X".

Write path (A sends a message to B)

  1. A holds a WebSocket to Chat Server 1, through the L4 load balancer. The JWT from the Auth service was validated when the socket opened, and Chat Server 1 registered A -> chat-server-1 in the Redis session store.
  2. A sends the message over the socket with a request_id (a UUID generated on the client) as an idempotency key. If the ack is lost and A retries with the same request_id, the server returns the original msg_id instead of storing a duplicate.
  3. Chat Server 1 forwards it to the Message Service, which assigns a Snowflake msg_id.
  4. The Message Service persists it to Cassandra, partitioned by conv_id and sorted by msg_id. Once the write succeeds, A gets a "sent" ack.
  5. The Message Service looks up B in Redis:
    • B online: deliver through B's chat server (Chat Server 2), and B acks back ("delivered").
    • B offline: publish to the offline queue (Kafka). The Push Notification Service consumes it and sends a notification through APNs / FCM.

Media doesn't go through the socket. A uploads the file over HTTPS (API Gateway -> Media Service -> S3) and the message carries only the media URL.

Read path

  • Live messages arrive pushed over the WebSocket; B doesn't poll.
  • Opening a conversation reads history from Cassandra: one partition (conv_id), newest first, paginated by msg_id:

    SELECT * FROM messages
    WHERE conv_id = ? AND msg_id < ?   -- cursor: oldest msg_id already shown
    ORDER BY msg_id DESC
    LIMIT 50;
    
  • Media is downloaded from the CDN using the URL in the message, not from the chat servers.

Offline sync

The push notification only tells B there is something new. The messages themselves come back when B reconnects:

  1. B opens a new WebSocket (possibly to a different chat server), the JWT is validated, and the session store is updated to point at the new server.
  2. B sends the last msg_id it received.
  3. The server replays everything newer from the Message DB. Since msg_id is time-ordered, this is a range scan: msg_id > last_received per conversation.
  4. B acks what it received, and the server advances B's sync cursor, so a reconnect halfway through just resumes from the last ack.

Delivery is at least once: a replay can resend a message B already has, and B drops it by msg_id.

Heartbeats over the socket keep presence (online / last_seen) fresh in Redis. When they stop, the session entry expires and B counts as offline, so new messages take the push path.

Pending questions

  • Why do we need an L4 load balancer? Can't we use L7, and what's the advantage of each?
  • Can WebSocket connections be relayed through L4 load balancers?
  • How does offline sync actually work? How does the server know which conversations to sync, and what if the user has thousands of chats?