
Somewhere around 2023, "prompt engineer" was a job title with a six-figure salary attached and a skill that mostly meant knowing a few magic phrases. By 2025, models had gotten good enough at guessing intent that the magic phrases stopped mattering, and a lot of people concluded the skill itself had evaporated with them.
It did not evaporate. It moved. The question worth asking in a hiring loop is no longer "does this person know clever phrasing tricks", it is "can this person tell a model exactly what it needs to solve a specific, real problem, and know when the model has stopped being useful." That is a durable skill, it is unevenly distributed across engineers who all claim to have it, and it is straightforward to see if you watch someone work instead of asking them to describe it.
Yes, and the strongest evidence is not a survey, it is what the model vendors themselves ship. Anthropic maintains a full prompt engineering guide with dedicated pages on context, examples, chain-of-thought and system prompts. OpenAI runs the same kind of guide for its own models. Neither company would maintain that documentation if the answer to "how do I phrase this" no longer changed the quality of what comes back.
What did happen is that the easy 20% got automated away. You no longer need to know that adding "let's think step by step" helps, because most current models do something like that by default. What is left is the harder 80%: supplying the context a model has no way to infer on its own, structuring a request so a multi-step or agentic task does not drift off course, and recognizing quickly when the tool has hit its ceiling on this particular problem. That is exactly the part a resume line cannot show and a multiple-choice quiz cannot grade.
Drop the idea that it is about wording. It is closer to three overlapping habits:
Specification. Turning a fuzzy goal into a request that contains the constraint, the relevant convention, the actual error text, and what a correct answer would need to look like. This is the same skill as writing a good ticket, aimed at a different reader.
Steering across turns. When the first answer is close but wrong, does the person add a new fact, or just repeat the same request with more emphasis? Across an agentic task with several steps, does the person check in enough to catch drift before it compounds, or let the model run to the end and then untangle what happened.
Knowing the edge of usefulness. Every model has a point on a given task where more prompting stops helping. Recognizing that point and switching to doing the work directly, or to a different tool, is arguably the most valuable half of the skill, and the one almost nobody was ever taught explicitly.
Those three habits are visible on a screen recording inside about ten minutes. They are not visible in a conversation about how someone approaches AI tools, because everyone gives roughly the same answer to that question, and the interesting variance lives in the doing rather than the describing.
A written prompt engineering quiz tests whether a candidate can name techniques: few-shot examples, role prompting, chain-of-thought. That is trivia now. It has been in every "top 10 prompt engineering tips" listicle for two years and it is exactly the kind of thing a candidate's own AI assistant would happily generate for them before the interview.
A take-home has a different failure. The candidate does the prompting on their own machine, over hours, with no time pressure and unlimited retries, then hands you a finished result. Everything interesting, how many turns it took, what they tried and threw away, whether they noticed the model was wrong before or after running it, happened off camera. You get the same artifact whether it took three prompts or thirty.
Both formats measure the wrong unit. The unit that matters is the sequence of requests and checks, not the final answer, because the final answer is reachable by very different paths that predict very different outcomes back on the job.
Three tasks, each aimed at one of the three habits above.
An underspecified change in a codebase the candidate has not seen. Give them a real repository with its own naming conventions and a vague ticket: "make search also match on tags." A generic prompt gets a generic, off-convention answer. Watch whether the first request already includes the file layout, the existing pattern for similar features, and what "match" should mean here, or whether the candidate fires a bare request and then spends ten minutes arguing with a plausible but wrong result.
A multi-step task with a planted trap. Ask for something that takes several actions: add a field, migrate existing data, update the two call sites that use it. Model agents left alone on multi-step tasks reliably miss one of the call sites or apply the migration in the wrong order. The candidate who checks in after each step catches it early. The candidate who lets the agent run to completion has to reconstruct what happened, which is a much harder job than steering would have been.
A task with a hard ceiling. Something the model genuinely cannot do well from this vantage point, for instance a bug that depends on state on the actual machine that no amount of rephrasing will surface. The candidate who keeps rewording the same request for fifteen minutes has told you something. The candidate who tries twice, concludes the model cannot see what it needs, and goes to read the logs directly has told you something more useful.
Three rows, and you can fill them in from a recording without arguing about it afterward.
| Habit | Weak | Strong |
|---|---|---|
| Specification | Bare request, generic answer accepted as-is | First request includes the real constraint, convention and error; one adjustment gets it close |
| Steering | Rewords the same ask, or lets a multi-step task run unchecked to the end | Adds a new fact each turn; checks in between agent steps and catches drift early |
| Edge of usefulness | Keeps prompting past the point of return | Notices the ceiling within two attempts and switches to direct work |
Score the behavior, not the final diff. Two candidates can land the same working code through very different processes, and only one of those processes will hold up on a problem you have not tested yet.
If you already run a broader AI-use interview, you may be wondering whether this is the same thing again. It is not, and it is worth being precise about the difference so you do not build two rounds that measure the same axis.
Four AI skills and how to test each covers "giving the model what it needs" as one of four behaviors, alongside choosing whether to use the tool at all, checking the output, and knowing what not to send it. That is the right scope for a general engineering hire, where prompting is one input among several. This post is the zoomed-in version, for when the request itself is the deliverable: agent instruction design, internal tool prompts, system prompts that other people will reuse, or any role where getting the request right is most of the job. If you are hiring a generalist backend engineer, run the broader interview. If you are hiring someone whose output is largely the prompts and specifications other engineers or agents will run against, run this one specifically, and score it on its own.
Everything above depends on seeing the process, not just the result, which means a screen and terminal recording, not a shared document after the fact.
In EasyEnv the candidate gets a real machine with the actual codebase, an AI assistant enabled, and the session recorded end to end. Reviewing afterward means watching what went into each request, how many turns it took to land, and whether they caught drift on the multi-step task before or after it had already gone wrong. That sequence is the whole assessment. The final commit, on its own, tells you almost nothing about how it was produced.
Tell the candidate up front that using the model is expected and will not be held against them. The interesting signal is what they do with it, not whether they reach for it.
Pull up your last three AI-assisted interviews and check one thing: could you say, from what you saw, whether the candidate's first prompt on the hardest task already contained the fact that mattered?
Run live coding sessions and take-home challenges in real production environments. Watch sessions back, score consistently, and hire with confidence.
More posts you might like
CoderPad is a strong, well-built live coding tool for teams with a steady interview pipeline and a recruiting function to run it. A ten-person startup doing its own hiring is a different buyer with different constraints. Here is what changes, and what to check in any CoderPad alternative for startups before you commit.
Every vendor calls its test a practical developer assessment and none of them define the word. Here is a four-question test you can run on any assessment in five minutes, worked through on a real example, to tell a work sample from a puzzle wearing a nicer outfit.
Read moreA live coding interview environment is more than a shared editor on a call. Here is what it actually needs to include, and how the three common setups compare when you look past the demo.
Read more