We measure AI literacy: how well a candidate works with AI, not whether they use it.
The candidate gets a real Linux box, a real task, the models you pick and a token budget. We record every prompt and every edit, then score how they steer the model, how they check its output and how much it cost them to get there.
Can they steer it?
Do they give the model context, constraints and a definition of done, or paste the task and hope?
Do they check it?
Do they read the diff, run the code and push back when the model is wrong, or accept whatever comes out?
Do they pick the right model?
A small, cheap model for boilerplate, a frontier model for the hard bug. Or the biggest model for everything.
Do they watch the cost?
Every interview has a token budget. We see whether they finish inside it or burn it on retries.
Each question with AI turned on gets a rating on four pillars, with a short note pointing at the prompt or the edit that earned it.
Understands AI limitations, hallucinations, bias, and reliability. Knows when the model is bluffing, when its answer is plausible-but-wrong, and which problems it cannot reach at all.
Writes clear, structured prompts to get high-quality results. Gives context, constraints, examples, and a definition of done so the model produces a real answer instead of a guess.
Verifies outputs, spots mistakes, and applies human judgment. Runs the code, reads the diff, checks the tests, and pushes back when the model is confidently wrong.
Uses AI securely, ethically, and without exposing sensitive data. Knows what to never paste into a model, respects licensing and attribution, and stays inside the team's guardrails.
Claude, ChatGPT, GLM, Gemini, DeepSeek, Qwen and the rest of OpenRouter, through one gateway. You choose which ones a role allows. No candidate needs their own API key.
We run on OpenRouter, so Mistral, Grok, Kimi, Llama and hundreds more are available too. If a model is out today, you can use it in your next interview.
Set a limit on the interview, for example “this interview can spend €5 on tokens”. The candidate sees what is left. When it runs out, the AI stops answering.
On the job, AI is not free. An engineer who sends the whole repo to the most expensive model on every retry costs you money every day. The report shows the spend, the tokens and which model each prompt went to, so you can tell who got the result cheaply and who just got it.
Opus for everything, including renaming variables. Re-sent the whole repo on every retry.
Budget ran out at minute 38. Task not finished.
Haiku for boilerplate and docs, Sonnet for the one hard bug. Sent only the files that mattered.
Finished at minute 31, tests green.
Pick the mode that matches the role. Score the candidate on how well they operate in it - and on whether they recognise which mode the task calls for.
Does not operate effectively in this mode at all.
Some capability but inconsistent. Needs heavy mentoring.
Reliably effective in this mode. Ready for the job.
Operates at a level you would want to learn from.
No spreadsheets, no honour system. The whole session is recorded, scored, and ready to replay.
Candidates get a fresh Linux VM with their stack and the models you allow. Same models and same budget for every candidate, so the comparison is fair.
Every prompt, every model response, every accept or reject is captured and timestamped against the code change. Reviewers see the collaboration, not just the artifact.
Tag each task as AI-Free, AI-Assisted, or AI-Directed. The rubric and the dashboard adapt, so you compare candidates on the same axis.
Run an AI-literacy interview on EasyEnv in under 10 minutes. Real box, real models, a real budget, recorded end-to-end.