
A distributed systems assessment is not a system design interview with harder vocabulary. A system design interview asks a candidate to describe a system that does not exist yet. A distributed systems assessment hands them one that already exists, already runs, and is about to do something wrong, and watches what they do about it. The two formats test different things, and most teams only run the first one.
That gap matters because the failure modes that define distributed systems work, a replica falling behind, a promotion racing a client that has not noticed the primary is gone, a retry landing twice, do not show up in a conversation about tradeoffs. They show up when a process actually dies. This post is about how to build an assessment around that moment: what to run, what to break, and how to score the fifteen minutes after it breaks.
Ask a candidate to explain the CAP theorem and most will get through it fine: under a network partition, you pick consistency or availability, you cannot have both. That answer is real knowledge and it is also, on its own, worthless as a hiring signal, because reciting the theorem and living inside its consequences are different skills. The engineer who has actually run a replicated system knows that "pick consistency" is not one decision made once. It is dozens of small decisions: does this specific read need the primary or can it tolerate a stale replica, does this write need to fail loudly if the quorum is gone or queue and retry, does the client back off or hammer the same dead endpoint for the next thirty seconds.
A distributed systems assessment has to surface those decisions one at a time, under a fault that is actually happening, not hypothesized. That means three things the format needs that a conversation does not:
Here is a concrete version of this we run: a workspace with a Postgres primary and a streaming hot standby, each on its own box, replicating over the network the candidate can see. In front of both sits a small connection pooler that routes writes to the primary and spreads reads across both nodes, so the application layer has one address and does not know or care which physical node answers a given query, which is exactly the illusion a distributed system is supposed to maintain.
Then we take the primary down mid-session. Not gracefully, not with a maintenance banner, just gone, the way a host actually dies. The brief is short: "the application is failing writes, find out why and get it healthy." Nothing in the prompt says the word "replica" or "promote." Finding that out is the assessment.
What we watch for in the first two minutes tells us more than the eventual fix does. A candidate who checks whether the pooler is still routing to a dead host, confirms the standby is caught up before touching anything, and only then promotes it, is showing the instinct that keeps a 2am incident from becoming a longer one. A candidate who immediately promotes the standby without checking replication lag first is optimizing for "traffic is flowing again" over "traffic is flowing again to a copy of the data we can trust," which is the exact tradeoff a real on-call engineer has to make correctly under pressure. Both candidates can end the session with a green health check. Only one of them made the decision you want on your team.
The promotion itself is one command, pg_ctlcluster promote on the standby, but the interesting part is everything around it: did they confirm the primary was actually unreachable rather than just slow, did they check the standby was not itself lagging before cutting over to it, did they update the pooler's config or leave it pointing at a host that no longer exists, did they think about what happens to the writes that were in flight when the primary died. None of that shows up if the interview stops at "and then we'd promote a replica."
Beyond primary failure, a handful of other faults consistently separate candidates who have operated distributed systems from candidates who have only read about them.
Replication lag under load. Generate write traffic, let the standby fall meaningfully behind, then ask the candidate why a read immediately after a write is returning stale data. The theorist says "eventual consistency" and stops. The operator asks which reads in this application actually need the primary, proposes read-your-writes for the one that matters, and leaves the rest on the lagging replica because paying for strong consistency everywhere is the wrong trade.
A retry that lands twice. Kill the connection right after a write commits but before the client gets the acknowledgment, so the client retries and the write happens again. Ask what the customer sees. A candidate who has been burned by this in production reaches for idempotency keys or a unique constraint on the operation, unprompted. A candidate who hasn't will describe retries as an unqualified good thing, which is the exact belief that produces a double-charged customer.
Quorum loss. With three or more nodes, cut two of them off from the third and ask what should happen to writes on each side. The right answer names the minority side refusing writes rather than accepting them and creating a fork to reconcile later. This is the single clearest place to watch whether "availability versus consistency" is a slogan the candidate has memorized or a decision they know how to make correctly when it costs them something.
A slow node mistaken for a dead one. Add latency, not a crash, to one node's health check response. Candidates who jump straight to "evict it from the cluster" without checking whether it is actually behind or just briefly slow will happily engineer a cluster that ejects its own healthy members under load, a failure mode that is more common in production than an outright crash.
The scenarios above are only useful if two interviewers watching the same session reach the same verdict, which is the same problem we cover in more general terms in how to write a technical interview rubric. For a distributed systems assessment specifically, score against the sequence of decisions, not the final state of the cluster:
A candidate who runs out of time with the cluster still degraded, but who narrated correct reasoning the whole way, is a stronger hire than one who restored green dashboards by luck and cannot explain why the fix worked. That distinction is exactly what a shared recording and a written scorecard capture and a verbal debrief afterward loses, which is the same argument we make in the practical interview scorecard.
If your team also runs a broader system design conversation before or after this, the two formats complement each other rather than compete: our guide to system design interview questions covers how to pick and score the discussion half. This post is about the fifteen minutes after the discussion, when the design has to survive an actual node going down.
You do not need a three-region, nine-node deployment to run this well. Two nodes and one real fault, injected mid-session instead of described in a prompt, already separates candidates more sharply than another round of "what is the CAP theorem." Pick one fault from the list above, write down what a strong response looks like before the candidate walks in, and see how many of your last five distributed systems hires would have caught the stale replica before promoting it.
For the deeper background on why partitions behave the way they do and how surprising the failure modes of real databases can be, Jepsen's analyses of distributed database consistency are worth reading before you write your first scenario; several of the faults above are drawn directly from failures Jepsen has documented in production systems.
Run live coding sessions and take-home challenges in real production environments. Watch sessions back, score consistently, and hire with confidence.
More posts you might like
Teams hire platform engineers on Kubernetes depth and Terraform fluency, then wonder why nobody notices the invoice climbing. Cost awareness is an engineering skill, it is easy to test on a running environment, and almost nobody tests it.
Under the EU AI Act, AI used to screen or evaluate candidates sits in the high risk category. Here is what that actually asks of you as an employer, the questions to put to your assessment vendor, and how to structure a process you could defend.
Read moreWhen the model has shell access, the question is no longer whether a candidate can write the code. It is whether they can supervise something that writes it faster than they can read it. Here is how to test that on a real repository.
Read more