
Most Node interviews have the same middle section. Explain the event loop. What is the difference between a microtask and a macrotask. What does process.nextTick do.
Candidates answer it well, because it is the most written-about topic in the ecosystem. Then the service they are hired to run restarts itself at 3am and nobody can say why.
The Node failures that cost money are quieter than the interview suggests. A promise nobody awaited. A stream that never closed. A retry loop with no backoff, hammering a dependency that was already struggling. A synchronous call in a request path that stops everything else for 400 milliseconds at a time.
None of that is a knowledge question. All of it is visible in an hour on a running box.
Because it is the part of Node that is easy to ask about, and because it feels fundamental. It is fundamental. It is also now the single most memorisable answer in backend interviewing, and a model will give it faster and more completely than any candidate.
What stays hard is reading a service that is misbehaving while you are responsible for it. That skill does not come out in conversation, because the conversation has no logs in it.
The setup. A service behind a process manager. Every few minutes it exits and comes back. In the logs there is a rejection that nobody handled, triggered by one specific input shape that arrives occasionally. Recent Node versions terminate the process on an unhandled rejection by default, which is documented behaviour rather than a mystery, and the crash loop looks like instability rather than a bug.
Ask: this restarts a few times an hour. Find out why and tell me what you would do about it.
Strong looks like: they find the rejection and trace it back to the call that produced it, rather than catching it globally and declaring victory. They can say which failures should crash a process and which should be handled, and they know the difference between making the symptom stop and making the request fail honestly.
Weak looks like: adding a global handler that swallows everything. The restarts stop, the bad input keeps being accepted, and the next person inherits a service that lies about its own health.
Why it works: it separates people who have operated Node from people who have only written it. Both groups can explain promises. Only one has watched a process manager paper over a bug for a month.
The setup. A service whose heap grows steadily under modest traffic. The cause is ordinary: an in-memory cache with no bound, or a stream that is never closed, or listeners attached per request to a long-lived emitter.
Ask: here is the graph. What is holding memory?
Strong looks like: they get evidence before theorising. A heap snapshot, or two taken a few minutes apart, and a look at what grew. They know the difference between a leak and a cache doing its job, and they ask what the expected working set is before calling anything a bug. They also check whether the growth tracks requests or time, which narrows it immediately.
Weak looks like: restarting on a schedule. It is a legitimate temporary measure, said out loud with a plan behind it. As a first answer, it tells you the rest of the diagnosis is not coming.
Why it works: memory is where Node's convenience quietly becomes a liability, and the diagnostic habit transfers to every other kind of production problem. We wrote about the same habit from the platform side in the observability skills assessment.
The setup. The service calls a dependency that is returning errors. The client retries immediately, three times, on every request. Under load that turns a struggling dependency into a dead one, and the logs are full of timeouts that look like the dependency's fault.
Ask: the downstream team says we are the reason they went down. Are they right?
Strong looks like: they read the retry policy before defending anybody. They know that immediate retries multiply load exactly when the system can least afford it, and they can name the fixes: backoff with jitter, a cap, a circuit breaker, and deciding which operations are safe to retry at all. The best candidates raise idempotency without being prompted, because retrying a write is not the same as retrying a read.
Weak looks like: treating retries as free, or removing them entirely to make the argument go away.
Why it works: it is a judgment question wearing a code question's clothes, and it is one of the most common ways a small outage becomes a large one.
The setup. One endpoint does something synchronous and expensive. A large JSON parse, a synchronous hash, an accidentally quadratic regular expression. Under concurrency, unrelated endpoints get slow, which sends everyone looking in the wrong place.
Ask: the health check times out when this report endpoint is used. Explain the connection.
Strong looks like: they connect single-threaded execution to the symptom without needing the hint, measure where the time goes, and propose something concrete: move the work off the request path, stream it, or make it asynchronous. They also say what they would monitor so this is visible next time rather than inferred.
Weak looks like: adding more instances. It helps, hides the cause, and costs money every month afterwards.
Why it works: this is the one place where the event loop knowledge everybody can recite actually has to be used, under evidence, on someone else's code.
Four rows, the same for every candidate:
None of those rows is "can explain the event loop". That shows up for free in task four, where it matters.
All four tasks need a service that is actually running: a process manager restarting it, a heap to snapshot, a dependency to overload, a load generator. A snippet in a document carries none of that.
EasyEnv gives the candidate a real Linux box with the service deployed, the unbounded cache already committed, the retry policy already wrong and the slow endpoint already in place. Every candidate gets the same recipe, so the bugs are identical and comparison is fair. The box is disposable, so shell access costs you nothing. The session is recorded, terminal and screen, so you can see whether they took a snapshot or guessed.
The limits, plainly. This is a 45 to 60 minute session and it belongs after a short screen. Reviewing it needs someone who has run Node in production, which is a smaller group than the people who can review a code sample. And for a role that is mostly front end with a little API work, two of these four tasks are more than the job needs.
Stop asking people to recite the event loop. Give them a service that restarts itself, a heap that climbs, a retry storm and a blocked handler, and watch what they open first.
If you only run one, run the retry storm. How someone reasons about the damage their own service does to its neighbours is the clearest seniority signal Node hiring has.
Run live coding sessions and take-home challenges in real production environments. Watch sessions back, score consistently, and hire with confidence.
More posts you might like
Java interviews still run on trivia that a study guide covers in an evening, while the expensive problems live in transaction boundaries, pool exhaustion and upgrades nobody dares start. Here are four tasks on a running Spring service, and what each one separates.
Most teams describe their AI adoption as "we're figuring it out," which is not a stage anyone can plan against. Here is a five-stage maturity model for AI literacy across an engineering team, how to tell which one you are actually in, and what moves you up one.
Read morePrompt engineering did not disappear when models got better, it changed shape. Here is what the skill actually looks like in 2026, why a resume line or a quiz cannot measure it, and the tasks and rubric that do.
Read more