๐ฏ The takeaway, first
Push for normal users, pull for celebrities โ run both. Precompute timelines for the 99.9% of users with small follower counts, and assemble celebrity posts at read time. Think of it like a newspaper: regular columnists get their article printed into every subscriber's paper overnight (push); the once-a-year royal wedding special gets inserted only when someone actually opens that page (pull). Anyone who says it's push or pull hasn't built one at scale.
๐ซ Misconception, busted
"You must choose push OR pull fan-out." Nearly every write-up frames this as a binary choice. In reality, every large feed runs a hybrid: push below a follower-count threshold, pull above it. The interesting interview discussion isn't which one โ it's where you draw the line and how the two paths merge at read time.
Requirements
Say these out loud before drawing a single box. Interviewers score this.
Functional
- Post (text โค 280 chars, optional media)
- Follow / unfollow users
- Home timeline, reverse-chronological, paginated
- View any user's own posts (profile timeline)
- Likes / reposts (mention, don't deep-dive unless asked)
Non-functional
- Timeline read p99 < 200ms
- 99.9% availability; reads must survive partial outages
- Eventual consistency is fine โ a post visible in seconds, not ms
- Scale: ~300M MAU; one celebrity post must not take down the system
- Media uploads shouldn't block posting
Back-of-the-envelope math
X-scale, circa 2020. Assumptions labeled โ interviewers care that you state them.
| What | Assumption | Math |
|---|---|---|
| Daily posts | 200M posts/day | โ 2,300 writes/s avg, ~5ร peak โ 12K/s |
| Timeline reads | 100M DAU ร 50 views | 5B/day โ โ 58K reads/s avg |
| Post storage | ~1 KB per post (text + metadata) | 200 GB/day โ ~73 TB/year before media |
| Median fan-out | median user โ 200 followers | 2,300 ร 200 = 460K timeline writes/s |
| Celebrity post | 1 user with 100M followers posts once | 100M writes from a single post โ more than a full day of normal traffic |
| Timeline cache | 300M users ร top 800 posts ร 16 B/post-id | โ 3.8 TB of hot Redis โ very shoppable, shard it |
Go deeper: why 800 posts per cached timeline?
A timeline page is ~20 posts; users rarely scroll past ~40 pages. Caching 800 covers the realistic scroll depth. Older posts fall back to a "load more" path that reads the post store directly. The cache is a window, not the archive โ this is what keeps the 3.8 TB number sane instead of ballooning to petabytes.
Architecture
The centerpiece. Write path fans out; read path merges.
flowchart LR
C[Mobile / Web client] --> G[API Gateway]
G --> PS[Post Service]
G --> TS[Timeline Service]
PS --> PDB[(Post Store
Postgres / Cassandra)]
PS --> M[Media Service]
M --> BS[(Blob Store
S3 + CDN)]
PS --> K[Fan-out Queue
Kafka]
K --> FW[Fan-out Workers]
FW --> GR[(Follow Graph
Postgres)]
FW -->|normal user| RC[(Timeline Cache
Redis Cluster
sorted set per user)]
FW -->|celebrity| CI[(Celebrity Index
recent posts of top ~10K accounts)]
TS --> RC
TS --> CI
TS --> PDB
TS --> C
Two write paths, one read path. A post from a normal user is fanned out by workers into each follower's Redis sorted set (timeline:{user_id}, score = timestamp). A post from a celebrity skips the fan-out entirely and lands in the Celebrity Index. At read time the Timeline Service pulls the precomputed timeline and recent posts from followed celebrities, merges them by timestamp, and returns one list. Media never touches the post path โ it uploads straight to the blob store and the post just carries media IDs.
Go deeper: how do you decide who's a "celebrity"?
A follower-count threshold (say, 100Kโ1M) recomputed by an offline job, plus a manual override list. Edge case: an account becoming famous mid-day. Handle it with a promotion job that stops their push fan-out and backfills the Celebrity Index โ their older posts stay in followers' cached timelines, new ones get pulled. Slight duplication at the merge step is fine; dedupe by post ID.
Component deep-dives
Posting as a normal user โ push fan-out
sequenceDiagram
participant C as Client
participant API as Post Service
participant DB as Post Store
participant K as Kafka fan-out topic
participant W as Fan-out Worker
participant G as Follow Graph
participant R as Redis Timelines
C->>API: POST /v1/posts {text}
API->>DB: INSERT post (durable first!)
API->>K: enqueue {post_id, author_id}
API-->>C: 201 Created (fast โ fan-out is async)
W->>K: consume job
W->>G: get follower IDs (paginated)
loop each follower batch
W->>R: ZADD timeline:{fid} ts post_id
W->>R: ZREMRANGEBYRANK (trim to 800)
end
The post is durable before the user gets a 201. Fan-out is async โ the user sees their own post instantly via a read-your-own-writes path, followers catch up within seconds.
Reading a timeline โ hybrid merge
sequenceDiagram
participant C as Client
participant T as Timeline Service
participant R as Redis Timelines
participant I as Celebrity Index
participant D as Post Store
C->>T: GET /v1/timeline?cursor=
par fetch both paths
T->>R: ZRANGE timeline:{me} (precomputed)
and
T->>I: recent posts from followed celebrities
end
T->>T: merge by timestamp, dedupe post IDs
T->>D: hydrate post bodies for IDs (MGET cache first)
T-->>C: 20 posts + next cursor
Go deeper: cursor pagination, not offsets
OFFSET 400 LIMIT 20 breaks the moment a new post lands โ items shift and users see duplicates or gaps. Use a cursor: ?before_ts=&before_id=. The client passes the oldest timestamp+ID it has; the server returns the next 20 older items. Stable under concurrent writes, and it composes cleanly with the merge step.
API + data model
POST /v1/posts { "text": "...", "media_ids": [] } โ 201 { "post_id": "p_9f2..." }
GET /v1/timeline?before_ts=&before_id=&limit=20
GET /v1/users/{id}/posts (profile timeline โ reads post store directly)
POST /v1/follow/{user_id} โ 204
DELETE /v1/follow/{user_id} โ 204
posts(post_id PK, author_id, text, media_ids[], created_at)
follows(follower_id, followee_id, created_at) -- PK (follower_id, followee_id)
Redis:
timeline:{user_id} โ ZSET, score = created_at_ms, member = post_id (trimmed to 800)
celeb:{user_id} โ ZSET of recent posts, only for top ~10K accounts
post:{post_id} โ HASH of hydrated post body (cache, short TTL)
Trade-offs
| Decision | Option A | Option B | Pick |
|---|---|---|---|
| Fan-out | Push: fast reads, write amplification | Pull: cheap writes, slow reads | Hybrid โ push under the celebrity threshold |
| Timeline store | Redis sorted sets (fast range queries) | Cassandra wide rows (cheaper, colder) | Redis for hot window; Cassandra as the archive |
| Fan-out delivery | At-least-once (dupes possible) | Exactly-once (slow, complex) | At-least-once + dedupe by post ID at merge |
| Consistency | Strong (followers see post instantly) | Eventual (seconds of lag OK) | Eventual โ nobody riots over a 3s delay |
Failure modes
What I'd actually build
Opinionated. Steal this for the interview.
Stack: Go API services (fast, boring, great JSON throughput) behind an ALB. Kafka for the fan-out queue with idempotent consumers. Redis Cluster sorted sets for timelines, Postgres (partitioned by month) for posts + follows with read replicas, S3 + CloudFront for media. Celebrity detection: nightly Spark/batch job over the follow graph + manual override list. Deploy on Kubernetes with HPA on the fan-out workers keyed to Kafka consumer lag. Pagination cursors everywhere; feature-flag the celebrity threshold so you can tune it live.
Why not Cassandra for posts? At 73 TB/year, Postgres with monthly partitions and read replicas is simpler to operate and plenty fast for point lookups by post ID. Reach for Cassandra when a single region's Postgres can't keep up โ that's a year-3 problem, not an interview problem.
Interview tips
- Open with the numbers. "200M posts/day, 58K timeline reads/s, and one celebrity post = 100M writes" frames every decision after it.
- Name the celebrity problem before they do. It's the question behind the question. Say "the naive push design dies on the first celebrity post" in minute five.
- Drive to hybrid yourself. Don't wait to be led โ propose the threshold and the merge step, then discuss tuning it.
- Mention read-your-own-writes. Async fan-out means the author must see their post instantly; handle it as a special case.
- Close with ranking. "Reverse-chron is v1; v2 is ranked โ that's a separate ML system reading the same stores." Shows scope control.
๐ฌ Interactive: fan-out simulator
Post as a normal user or a celebrity, under push, pull, or hybrid โ and watch the write amplification happen.
Analogy check: push = printing the article into every subscriber's paper overnight. Pull = inserting it only when someone opens that page. Hybrid = the newsroom picks per author.