Practical preparation / architecture

System design interviews: requirements, trade-offs and a worked service

A practical system design playbook with a worked notification-service example, capacity reasoning, failure testing and a repeatable answer structure.

First question
What must the system do?
Core trade-off
Failure versus cost
Strong ending
Test the assumptions
Engineer sketching a resilient distributed service architecture
System design: an editorial illustration for the preparation path below.

Preparation map

Four moves before the interview

  1. 1

    Set the contract

    Clarify users, operations, latency, durability, scale and out-of-scope cases.

  2. 2

    Draw the first version

    Define API, data model, source of truth and request path.

  3. 3

    Stress the design

    Find bottlenecks, failures, retries and consistency choices.

  4. 4

    Validate the plan

    Choose metrics, rollout gates and recovery drills.

Visual system design path from requirements through data model, reliability, scaling and verification
The order prevents premature architecture: first define the contract, then prove where complexity is necessary.

01 / FIELD NOTES

What the interview is trying to learn

Microsoft names distributed systems, resilience, availability, replication and partitioning among possible topics. Amazon emphasizes practicality, accuracy, efficiency, reliability, optimization and scalability. Neither source says a memorized diagram is enough.

  • A design question is underspecified on purpose. Ask whether the priority is latency, durability, low cost, privacy, global reach or ease of operation. Say which requirement you are treating as the first constraint.
  • Translate broad language into checkable terms: expected reads and writes, peak factor, payload size, retention, recovery objective and acceptable stale data. Use round numbers as explicit assumptions, not claims about a real company.
  • A sensible first architecture often has an API, service, durable store and asynchronous worker. Add caches, queues, replicas and partitioning only when a stated requirement explains them.
  • Interviewers can change a constraint midway. Show which component changes and which invariants remain, rather than rebuilding the entire diagram from memory.

02 / FIELD NOTES

Worked example: a notification service

Suppose a product must send transactional email and in-app notifications. Users should see status, repeated requests must not produce duplicate sends and a temporary provider outage must not lose accepted work.

  • Contract: POST a notification with tenant, recipient, template, payload and idempotency key; GET its state. Confirm whether exactly-once delivery is required or whether at-least-once processing with deduplication is acceptable.
  • First design: authenticate and authorize the request; persist an accepted record with a unique idempotency key; enqueue a job; let a worker call the external provider; record attempts and provider outcomes.
  • Failure path: if a worker crashes after sending but before recording success, a retry may send twice. Prefer provider idempotency where available and make the user-visible state honest about uncertainty.
  • Scale path: partition queued work by tenant or destination only when traffic and fairness need it. Bound retries with backoff and a dead-letter queue, and protect one tenant from exhausting workers for others.
  • Verification: track accepted-to-sent latency, retry rate, duplicate suppression, queue age and delivery status mismatch. Roll out to a small cohort and rehearse provider outage and replay.

03 / FIELD NOTES

Choose the right depth for your level

  • Junior or mid-level: focus on a coherent API, data model, correct failure path and clear assumptions. It is better to explain one queue correctly than to name five infrastructure products.
  • Senior: show how you choose reliability targets, bound blast radius, handle migration and sequence a rollout. Explain who operates the system and what signal triggers rollback.
  • For any level, discuss privacy and security at the relevant boundaries: tenant authorization, sensitive payload retention, encryption, provider access and audit events.
  • If a component is unfamiliar, describe the property you need rather than inventing vendor behavior. State what you would test before adopting it.

04 / FIELD NOTES

A 40-minute rehearsal you can repeat

  • Minutes 0–7: clarify users, operations, scope and quality targets. Minutes 7–15: draw API, data flow and source of truth. Keep the first design intentionally small.
  • Minutes 15–25: walk through one read and one write, then inject two failures. State what the user observes and what the operator can recover.
  • Minutes 25–35: estimate the first bottleneck from your assumptions and add one justified scaling mechanism. Explain its new cost or consistency trade-off.
  • Minutes 35–40: name metrics, tests, rollout stages and a rollback trigger. Ask the interviewer which area deserves a deeper dive.

Original practice bank

Three tasks to work through aloud

These are ApplyDjinn practice prompts, not questions obtained from an employer. Use them to rehearse the reasoning and verification expected for the role.

Exercise 1

Design: rate-limited webhook delivery

Prompt: A service must deliver webhooks to customer endpoints that may fail or rate-limit requests.

Approach: Define per-tenant ordering and retry guarantees, persist events before enqueueing, use bounded backoff and idempotency, and expose delivery attempts to customers.

Test: Simulate endpoint 429, timeout after success, poison payload and a tenant whose endpoint remains down for days.

Follow-up: How do you prevent one failing tenant from blocking the queue?

Exercise 2

Design: read-heavy activity feed

Prompt: Users read a feed far more often than they create activities. Some users have millions of followers.

Approach: State freshness target and access control; compare fan-out on write versus read and consider a hybrid for very large accounts. Define the source of truth and cache invalidation boundary.

Test: Examine a deleted activity, privacy change, celebrity post and regional cache failure.

Follow-up: Which metric would tell you to change the fan-out strategy?

Exercise 3

Failure drill: queue replay

Prompt: A deployment accidentally replays a day of accepted jobs. What should happen?

Approach: Use stable event IDs, idempotent state transitions, provider dedupe where possible and a pause/reconcile plan before resuming sends.

Test: Distinguish data-plane duplicate suppression from audit records of the replay itself.

Follow-up: How would you prove the replay did not silently drop legitimate new jobs?

Source check

Primary sources

Company process details come from these official pages. The exercises and preparation framework are original editorial material; confirm your own interview format with the recruiter.

Continue preparing with evidence

Choose a current role, map its requirements to projects you can defend, then practice the round your recruiter actually scheduled.