
You have thirty minutes with a candidate, a job spec full of words you had to look up, and a hiring manager who will complain either way. Send too many people through and you are wasting engineering time. Send too few and the role stays open and that is your fault too.
Nobody trained you for the technical part, because there is no training that turns a thirty-minute call into a code review. That is fine. The recruiter screen was never supposed to be a code review.
This is about doing the part you can genuinely do well, doing it in a way engineers respect, and handing over something better than a gut feeling.
Three things, and only three:
That is it. You are not deciding whether they can code. Someone else does that with a real task. Your job is to make sure the person who does that step is not wasting their time.
The trick is asking things where the shape of the answer tells you what you need, not the content.
The best question you have, and it costs nothing to judge.
Good sounds like: they start with why it mattered or what was broken, then what they did, then how they knew it worked. They notice when they use a term you would not know and explain it without being asked. You come away able to repeat the story.
Weak sounds like: a wall of tool names with no story. Or an answer so vague it could be about any project at any company. If you cannot repeat back what they did, that is data. Engineers who cannot explain their work to you usually cannot explain it in a design review either.
Not a red flag: an accent, nerves, or slow speech. You are listening for structure, not polish.
Good sounds like: a specific dead end. "We assumed it was the database, spent two days there, and it turned out to be the load balancer." Real work has wrong turns in it.
Weak sounds like: nothing was hard, or the hard part was somebody else. People who have actually done the work remember the parts that hurt.
Good sounds like: clean edges. "I did the API, another engineer did the front end, I was not involved in the deploy pipeline." Comfortable saying where their work stopped.
Weak sounds like: "we" for everything, and no answer when you ask which part was theirs. This is the single most useful signal for spotting a CV that has been generously edited, and you do not need to know what an API is to hear it.
Good sounds like: a clear preference that matches the job, or an honest mismatch. "I have done four years of support work and I want to move to building." That is useful either way.
Weak sounds like: whatever they think you want to hear. Watch for the candidate who agrees with every part of the spec.
The single biggest upgrade to a recruiter screen is not a better question. It is passing on something the hiring manager can look at instead of your summary.
A summary says "strong communicator, five years Python, seems solid." The hiring manager cannot check that, so they redo your screen themselves. Now you have spent an hour of their time and they still think the shortlist is unreliable.
An artifact is different. A short piece of real work, done in a real environment, that the manager can look at for five minutes and form their own view: what the person did, what they tried, where they got stuck.
That does three things at once. It stops the manager repeating your call. It gives you cover, because your judgment is no longer the only thing standing behind the shortlist. And it moves the technical decision to the person whose decision it is.
Say this out loud to your hiring manager, and get agreement in writing:
That line keeps moving because nobody ever draws it. Draw it. A shortlist that gets blamed is usually a shortlist where the brief was "find good engineers" and nothing more precise was ever agreed.
If you are newer to this and want the broader version, we wrote how to become a technical recruiter engineers respect and breaking into technical recruiting without a tech background.
Most assessment tools output a score. A score does not help you, because you cannot defend it. If the manager asks "why is 78 good?", you have nothing.
EasyEnv outputs a session. The candidate works on a real Linux machine on a small task from the actual job, and the whole thing is recorded, terminal and screen. You do not have to understand what you are looking at. You attach it, and the hiring manager watches four minutes of it and knows more than your notes could have told them.
For agencies this matters twice over, because the artifact travels with the candidate to the client. Written answers get an AI-graded first pass, which gives you a plain-language summary to read even when the code is not something you can judge.
What it does not do: replace your call. Somebody technical still has to look. If your client wants a number they can sort a spreadsheet by, an auto-scored test does that and this does not.
Screen for reality, relevance, and whether they can explain their own work to you. Those three you can judge better than anyone. Technical ability is not on your list, and pretending it is gets you blamed for both outcomes.
Then hand over something the manager can look at. Your judgment stops being the only thing holding up the shortlist, which is the actual fix for the argument you keep having.
Run live coding sessions and take-home challenges in real production environments. Watch sessions back, score consistently, and hire with confidence.
More posts you might like
Generated code fails in a different pattern from human code, which is why normal review habits miss it. Here is the review order that catches the most in the least time, and the hiring task built from it.
Every candidate now claims AI fluency, so the claim carries no information. Here is what tool choice, context giving, verification and responsible use look like when you actually watch someone work, and the tasks that surface each one.
Read moreThe phrase went into the job ad and nobody can say what a candidate has to show to meet it. Here is AI literacy broken into four graded behaviors with level descriptors you can screen, interview and train against.
Read more