Why technical hiring should test how developers work with AI agents
Developers build with AI agents every day, then interview without them. Here is what a coding assessment should measure instead, and how kodwai scores it.

Most developers now build with an AI coding agent open next to them. They describe the change, read what comes back, run the tests, push back when it is wrong, and ship.
Then they interview, and the agent is gone. The loop asks them to reverse a linked list on a whiteboard, with no tools, in a setting that looks nothing like their job.
That gap is the reason kodwai exists.
What a typical interview loop actually tests
Strip away the details and most loops look roughly the same: a resume screen, a phone screen with knowledge questions, a take-home or an algorithm round, and an on-site with more of the same.
Each step tests something real, but mostly the same few things: recall, speed under pressure and practice at interviewing. None of them shows how a candidate works when the tools they use every day are in the room. Teams end up hiring people who are good at interviews, and passing on some who are good at building software.
How the work changed

| Before agents | With agents |
|---|---|
| Write most lines by hand | Describe the change, review the diff, iterate |
| Search the docs and forums | Ask with the codebase as context |
| Hours of boilerplate | Scaffold quickly, spend the time on design |
| Debug alone | Debug with a second pair of hands that needs direction |
What matters now is less about memorizing an algorithm and more about directing an agent toward a working, tested solution: breaking the problem down, saying precisely what you want, checking what comes back and recovering when it goes wrong.
Enter kodwai
The premise is simple: if developers work with AI agents, assess them working with AI agents.
kodwai is a platform where developers solve real coding challenges on their own machine with their own AI coding agent (Claude Code, Cursor, or Codex) and get scored on how well they direct the agent, across three axes: Direction, Outcome, and Lift.

How it works
1. Pick a challenge. The challenges are ticket-sized, real-world problems, not puzzles. Browse them at kodwai.com/challenges. The starter is Bookshelf REST API, a backend challenge with a 60-minute timer.
2. Solve it locally, with your own agent. Open a terminal wherever you want the project to live and run:
npx @kodwai/cli@latest challenge bookshelf-rest-api
The CLI signs you in, asks which agent you are using (Claude Code, Cursor or Codex), downloads the problem statement, starter files and tests, sets up a git repo and starts the timer. There is no browser sandbox. You work in your own editor, on your own machine.
3. Submit with one command.
npx @kodwai/cli@latest submit
That packages your code, git history, test runs, the agent transcript and the time taken.
4. Get your score. You get a score from 0 to 100 across three axes:

- Direction: how you steer, verify, and decompose.
- Outcome: what shipped, replayed and stress-tested to prove it holds.
- Lift: how far you beat a solo AI, not just that you passed.
Direction carries the most weight. The score is dominated by how you direct the agent, the part a careless prompt cannot fake. Passing tests is necessary but not sufficient. Every signal cites its own evidence from your transcript, commits and test runs, and your score comes with a confidence interval instead of false precision. More on the score itself is on the AI Collaboration Score page.
Your results go on a public leaderboard and build a profile you can point to.
Why this matters for developers
A whiteboard round rewards a narrow skill: solving puzzles under pressure without tools. It says little about how you work on a real codebase with the tools you actually use.
kodwai looks at the things that do show up at work:
- Decomposition: do you break the problem into pieces before handing work to the agent?
- Spec precision: do you say exactly what you want, with constraints and a definition of done?
- Verification rigor: do you run the tests and read the diff, or accept the first thing that compiles?
- Recovery: when the agent goes off course, do you diagnose and redirect?
- Outcome: does the result actually work, and is it code someone could maintain?
Solving challenges, the score, the profile and the leaderboard are free for developers. Every account gets 3 submissions scored on our key. After that, you connect your own Anthropic API key in Settings for unlimited submissions.
Why this matters for companies
Hiring managers want engineers who can work well with AI, and most interview loops cannot see that skill at all.
The kodwai hiring track runs the same idea as a private interview. You send a custom interview project with a one-link email invite, and the candidate solves it locally in Claude Code. Interviewers can watch the session live: the prompts, commits, test runs, tools used, time and API cost. kodwai then scores the candidate against your team's own per-role rubric, each dimension with a written justification, strengths and weaknesses. Your team adds its own scores and comments and can compare candidates side by side.
What it does not do yet: there are no ATS integrations, no published pricing and no video playback of the session. Not yet, and we will not pretend otherwise. Details are on the kodwai hiring track, and if you are designing the interview itself, I wrote a separate guide on how to interview developers who work with AI agents.
The bottom line
Technical hiring should test the job as it is done today: building real things with real tools, including the agent. kodwai lets developers show that, and lets teams see it.
To try it, create a free account and run your first challenge:
npx @kodwai/cli@latest challenge bookshelf-rest-api
try it yourself
Reading about it is the easy part.