
Ask a Java candidate to explain the difference between an interface and an abstract class. They will get it right. Ask how a HashMap resolves a collision and most will manage that too, because it is on the first page of every interview guide and has been for twenty years.
Now think about the last production incident your team had with a Java service. It was probably not caused by anyone misunderstanding a HashMap. It was a transaction boundary one layer too high, a connection pool sized for a laptop, a lazy load that worked in a test and threw in a controller, or a framework upgrade that nobody started because nobody could predict what would break.
Those are the things to interview for. None of them fit in a quiz, and all of them fit in a terminal.
Because it grades itself. Twelve questions, count the right answers, everyone leaves feeling the process was rigorous. It also feels safe: nobody gets blamed for asking the question that was asked at every other company.
The problem is that this knowledge is the cheapest thing on the list. Java's syntax and collections behaviour are revisable in an evening and answerable by a model in a second. What takes years is the judgment to look at a stack trace from a service that is failing under load and know which of the four plausible causes to check first.
So put the candidate in front of a running service.
The setup. A Spring service with a method that writes two rows and then calls a helper that writes a third. @Transactional is in a place that does not do what the author assumed, for example on a private method, or invoked from inside the same class so the proxy never applies (the Spring documentation on self-invocation is explicit about this). Somewhere in the data is a half-written record from last week.
Ask: this record should never have been able to exist. Explain how it did.
Strong looks like: they establish where the transaction actually begins before theorising. They know proxy-based annotations do nothing on a self-invoked call, and they can prove it rather than assert it, by logging, by a test, or by reading the configuration. They then say what they would change and what it costs, because moving the boundary outward is not free either.
Weak looks like: adding @Transactional to more methods until the symptom stops. The annotation is cheap to scatter and the understanding is the whole point.
Why it works: half-written data is the most expensive class of bug in a business system and the hardest to reason about afterwards. Someone who has been burned by it talks about it differently.
The setup. The service is fine with one user and falls apart with fifty. The connection pool is small, a slow downstream call holds connections while it waits, and the thread pool is sized the way it was on somebody's laptop. Logs show timeouts that point at the database, which is not actually the problem.
Ask: here is the load test and here are the logs. Tell me what is happening.
Strong looks like: they separate symptom from cause. They notice that the database is not busy, that the waiting is happening before a query is ever issued, and that the timeouts cluster the moment concurrency crosses a number. They talk about pool sizing as a consequence of latency and concurrency rather than a value you copy from a blog post, and they ask what the downstream service's own limits are.
Weak looks like: raising every number. Bigger pool, more threads, longer timeout. It makes the test pass and moves the failure somewhere quieter and later.
Why it works: the misleading-evidence shape here is what production actually feels like. We wrote about the same habit in a different seat in the database performance interview.
The setup. A service a few versions behind, with a change that touches everybody: a major Spring Boot jump, the javax to jakarta namespace move, or a JDK bump that turns a warning into an error. The build does not pass. Tests exist and some of them fail for unrelated-looking reasons.
Ask: get this building and tell me the order you would do the real upgrade in.
Strong looks like: they read the release notes before touching anything. They work in layers: compile first, then tests, then runtime behaviour. They can tell which failures are the upgrade and which were already broken. Crucially they talk about sequencing and risk: what ships on its own, what has to go together, what is reversible.
Weak looks like: fixing errors one at a time in the order the compiler emits them, with no sense of which ones are the same root cause, and no plan beyond green.
Why it works: most Java estates are old. The person you hire will spend real months on upgrades, and almost nobody is interviewed for it. This is also a good senior signal, because planning the order is the part that cannot be looked up.
The setup. A service with a slow leak, a heap dump or a flight recording already captured, and a graph that climbs across restarts.
Ask: what is holding memory, and what do you want to do about it?
Strong looks like: they open the dump rather than guess. They look for what is retaining objects, not just what is numerous. They can distinguish a leak from normal caching, and they ask what the expected working set is before calling anything a bug. If they do not know the tool well, they say so and still reason their way there, which is a better signal than fluent hand-waving.
Weak looks like: suggesting a bigger heap, or a GC tuning flag copied from somewhere, with no evidence about what is being retained.
Why it works: it is the clearest separator between someone who has operated a JVM in production and someone who has only written code that runs on one.
Four rows, the same for every candidate:
Notice that none of those rows is "knows Java". Fluency shows up for free while they work, in the same way an accent shows up in conversation.
These tasks need a running service, a database with the bad row already in it, a load generator and a heap dump that is real. A code snippet in a document cannot carry any of that.
EasyEnv gives the candidate a real Linux box with the service built, the misplaced annotation already committed, the pool already undersized and the dump already captured. Every candidate starts from the same recipe, so the problems are identical and comparison between people is fair. The box is disposable, so full access is not a risk. The session is recorded, terminal and screen, so you can see whether they read the release notes or started changing imports.
The limits, plainly. A session like this runs 45 to 60 minutes and needs an interviewer who knows the JVM to review it. Building the broken service the first time is real work, though you only do it once. And if the role is greenfield with no legacy estate behind it, task three is not the right use of the hour.
Stop asking how a HashMap resolves a collision. Put a transaction that silently did nothing, a pool that exhausts under load, an upgrade nobody started and a heap that only grows in front of the candidate, and watch what they open first.
If you only run one, run the upgrade. Whether someone can sequence a change by risk, rather than by compiler output, is the difference between a Java estate that stays current and one that quietly becomes a rewrite.
Run live coding sessions and take-home challenges in real production environments. Watch sessions back, score consistently, and hire with confidence.
More posts you might like
Most teams describe their AI adoption as "we're figuring it out," which is not a stage anyone can plan against. Here is a five-stage maturity model for AI literacy across an engineering team, how to tell which one you are actually in, and what moves you up one.
Prompt engineering did not disappear when models got better, it changed shape. Here is what the skill actually looks like in 2026, why a resume line or a quiz cannot measure it, and the tasks and rubric that do.
Read moreCoderPad is a strong, well-built live coding tool for teams with a steady interview pipeline and a recruiting function to run it. A ten-person startup doing its own hiring is a different buyer with different constraints. Here is what changes, and what to check in any CoderPad alternative for startups before you commit.
Read more