Think you're good with AI? Prove it on the leaderboard.
One hour, a real AI agent, and a broken app to rescue. The best scores fix it fast and clean. Get a score out of 100, a tier, and your rank against everyone else.
Free to play. Runs in your browser. Sign in to save and share your rank.
App live in 11m 40s. Clean root-cause fix, one wasted turn.
The challenge
Rescue the crashing API
An orders API 500s on an id that does not exist. This is the actual console you get. Drive the agent to a fix, then work out whether the fix it handed you is the right one: the tests you can see are only half the score.
What we score
It rates you, not the AI
Everyone gets the same agent. The score is your judgement: the result, how you delegate, and whether you catch its mistakes.
Get the result
Did it actually work? Outcome is graded first, because a run that does not ship is not a win.
Delegate well
Clear prompts, few wasted turns. The score rewards briefing the agent precisely and keeping it on the real problem.
Catch its mistakes
The agent will be confidently, subtly wrong. We reward the moments you catch it and steer it back instead of shipping the bug.
The leaderboard
Where will you land?
Every run is judged on the same challenge, so the score is comparable. Climb from Learning to Elite, beat your friends, and share a link that shows exactly where you rank.
Claim your rankHow it works
Three steps, one hour, one rank
Boot the broken stack
A real Linux machine spins up in your browser with a dead Django app, an editor, a terminal, and an AI agent. Nothing to install.
Drive the agent
You have an hour. Brief the agent, read its output, and steer it toward the real root cause. You direct, it types.
Get scored and ranked
When the app is live, the judge scores your run out of 100, hands you a tier, and drops you on the leaderboard.
Find out how good you really are
One broken app, one hour, one shareable rank. No setup, no install. Start in your browser right now.
FAQ
Questions, answered
It is a hands-on challenge that scores how well you work with AI, not how good the AI is. You get a real Linux machine with a broken app and an AI agent. You have an hour to drive the agent to find the real root cause and get the app live. We score the outcome, how cleanly you delegate, and whether you catch the agent when it is wrong.
You. The same agent is available to everyone, so the score reflects your judgement: how clearly you brief it, how few wasted turns you take, and whether you spot the moments it is subtly wrong. Two people with the identical model get very different scores.
Rescue a broken Django deploy. docker compose brings the container up, then it crashes on boot. The failure has a single real root cause. Drive the agent to diagnose it and get the app serving on http://localhost:8000. The best scores fix it fast and clean.
You direct the agent in plain language, so you do not have to write the fix yourself. Basic developer literacy helps you read the agent's output and catch its mistakes, but the test rewards how you delegate and verify, not how fast you type.
You have up to an hour, but the highest scores fix it fast and clean. It is free to play and runs entirely in your browser. Sign in to save your result and share it.
A shareable score out of 100, a tier, and your rank against everyone else who has taken it. Send the link to a friend and challenge them to beat it.
