An in-house team that interviews a weak candidate loses an hour. An agency that submits one loses standing with the client, and standing is the only thing being sold.
That asymmetry is the entire argument for agencies assessing before submission, and it is why the tooling that works for an internal hiring team often does not survive contact with agency work. The internal team optimises for not hiring the wrong person. You are optimising for the client saying yes to the next CV you send, which is a different problem with a shorter feedback loop and a much harder economic edge.
What you are really being paid for
Fill rate pays the invoice. Submission-to-interview ratio decides whether there is another requisition.
A client who interviews four of your five submissions stops looking at other agencies, because you have taken work off their plate. A client who interviews one in five has learned that reading your submissions is work, and the next role goes to three agencies at once. Nothing in your outreach fixes that. It is decided by whether the people you send can do the thing.
The trouble is that the artifact you screen on cannot tell you. A CV lists what somebody was near. Six years of Python at a company you have heard of is compatible with writing the service and with attending its standups.
Why the usual screening does not close the gap
The phone screen measures fluency about work, not the work. Some engineers explain their systems beautifully and some of the best ones do not, and the correlation between the two is weaker than a fifteen-minute call makes it feel. You are largely measuring how often somebody has been interviewed.
Client-side tests are too late. By the time the client runs their assessment, you have spent the submission. If it goes badly, the cost has already landed.
Generic coding tests measure the wrong thing and cost you candidates. A puzzle round tests puzzle practice. Worse, senior engineers with options decline them, so the filter removes exactly the candidates who make your ratio, and you never see who dropped out.
What agency work needs that in-house tooling does not
It has to run before submission and take under an hour. Your leverage exists in the window between sourcing and sending. An assessment that takes a candidate three evenings does not fit in that window and will not be completed by anybody currently employed.
It has to produce something you can forward. This is the requirement in-house platforms mostly ignore, because internally the assessor and the decision-maker are the same team. You need an artifact a client engineering manager will read and believe: what the candidate did, on what, with the evidence attached. A percentile score is not that. It arrives as your opinion in a different font.
It has to survive not having an engineer on staff. Most agencies do not have someone who can read a diff. The platform has to encode the judgment in the scenario and the record so a recruiter can report accurately what happened without pretending to evaluate the code. See screening engineers when you cannot read code for how that works in practice.
It has to be defensible when the client disagrees. Occasionally a client rejects someone who passed. If your evidence is a score, that conversation ends with a shrug. If it is a recording of the person doing the work, it ends with the client telling you what they weigh differently, which is worth more than the placement.
Candidates must not resent it. Your candidates are not applicants to you, they are your inventory, and you want them back. A round that feels like a real conversation about real work is one they tolerate. A timed puzzle is one they remember.
What the assessment should actually be
For agency screening the useful shape is a short, real, unglamorous task on a working environment. Not a puzzle, not a project. Something like: this service returns 500 on one endpoint, the environment is running, find out why.
That surfaces the thing your client cares about, which is whether the person can operate in a codebase they did not write. It is short enough to schedule, concrete enough that a non-engineer can follow what happened, and specific enough to the role that it does not feel like a hurdle.
In EasyEnv the candidate gets a real machine with the failing service on it, the session is recorded, and the output is a record of the work rather than a grade. You forward the record. The client watches four minutes of it at double speed and interviews the person. That is the whole loop, and it is short on purpose: every extra step is candidates lost.
Introducing it without losing your pipeline
Position it as preparation, not a hurdle. Candidates accept a round that is framed as making their submission stronger, especially when you tell them the client will see how they work rather than a CV.
Never run it on candidates you are not submitting. The assessment is not a sourcing filter. Use it on the shortlist you already believe in, or you will burn goodwill across the market for candidates you were never sending.
Tell the client it exists, before they need it. "Every submission comes with a recorded session" is a differentiator during the sale and an explanation when a rejection happens.
Measure one number. Submission-to-interview ratio, before and after, on the same client. If it does not move within two months, the assessment is wrong and you should change the scenario rather than the tooling.
The uncomfortable part
Assessing before submission means finding out that some of the people you were about to send cannot do the job. That is a smaller pipeline this month and a larger one next quarter, and the first month is genuinely worse. Agencies that stop at that point conclude assessment does not work for them.
It works. It just moves a cost that was previously invisible, paid in client relationships, into a cost that is visible and paid in submissions. Only one of those two compounds in your favour.
When the applicant pile is bigger than your team can read covers the volume end of the same problem.
What is your submission-to-interview ratio with your best client right now, and do you know it to one decimal place?