for hiring teams

See how candidates
really build.

Kodwai runs a real ticket as a private interview and shows you the whole session: the prompts, the recovery, the verification, the result. You watch it live, score it against your own rubric, and compare candidates on one board.

//live session
Jamie BrooksSenior Backend
activetime18:24cost$0.42tokens31.2k
~/kodwai/rate-limiter · zshREC 18:24
signals
time18:24cost$0.42tools9tests9 / 9
//01the old interview fails

The take-home is broken.

It is homework, not work. You wait days for a diff that hides the process, and the candidate with the most free time wins. You never see how they actually drive an agent.

// why the take-home fails

  1. It is homework, not the job.

    Graded blind, days later, on a final diff that hides every decision that mattered.

  2. It rewards free time, not skill.

    The candidate who can spend a weekend on it wins. That is not the signal you want.

  3. It is easy to outsource and easy to game.

    You cannot tell who actually drove the work, or how.

  4. You never watch them work.

    The one thing that predicts the job, how they steer and recover with an agent, is exactly what you never see.

//02see the real work

See exactly how candidates work.

The same engine as the developer challenges, pointed at hiring. You watch the process, not just the result, and you score it on terms you set.

liveclaude code · interview
18:24$0.42
you ›per-key sliding window, monotonic clock, no leak
claude ·scaffolds limiter.py, adds RateLimiter + store
$ pytest -q
✗ 1 failed · burst window races
claude ·adds per-key lock, scopes critical section
$ pytest -k concurrency
✓ 4 passed
+ commit: handle burst window + per-key lock
you ›now evict idle keys so memory stays flat
claude ·adds a sweep on the monotonic clock tick
$ pytest -q tests/test_eviction.py
✓ 6 passed
you ›what is the p99 under one hot key?
claude ·benchmarks it, 9.4k req/s, p99 0.21ms
$ ruff check limiter.py
✓ all checks passed
claude ·writes the PR body with the trade-offs
+ commit: evict idle keys, add benchmark
you ›
tools7
tests4/5
diff+38 -6

Watch the session live

You follow the whole session as it happens, every prompt, every tool call, every test run, every commit, with elapsed time and running API cost. It is a live dashboard, not a recording you wait for at the end.

RUBRICedit
Problem decomposition30
Direction & verification30
Code quality25
Recovery15
weights= 100%

Your problem, your rubric

You author the ticket, then define the dimensions and weights you score on. The AI grades every candidate 0-10 against exactly those terms.

8.4
/ 10vs your rubric
correctness9.0
code quality8.0
verification7.5
// justified per dimension"verified every claim before shipping."

A transparent AI score

Every finished session is scored against your rubric, 0 to 10, with a written justification for each dimension. Reasoned, not a black box.

AI review
8.4/ 10

Reverted a failing migration, then shipped a clean fix with tests.

your team
Maya, staff eng

Agree, strong recovery. I would hire.

add your score

Score it yourself too

Your team adds their own scores and comments right next to the AI score, on the same 0 to 10 rubric. The AI verdict and your read sit side by side, never one instead of the other.

candidate
decomp
direction
quality
overall
Jamie
8.6
8.1
8.5
8.4
Sarah
7.8
7.2
7.6
7.5
Kenji
7.1
6.4
6.8
6.8
scored 0 to 108+6.5 to 8under 6.5

Compare side by side

Put everyone for a role on one board and read the rubric as a heatmap. The same dimensions apply to every candidate, so the strongest and weakest work surfaces at a glance before you open a single transcript.

time limit60:00
budget$5.00
toolsallowblock
agentClaudeClaude Code
applied to every candidate

Time, budget, tools

Set the clock, cap the API spend, and allow or block tools per interview. The same constraints for every candidate, so the comparison is fair.

//03the whole session

The whole session, not the diff.

A take-home shows you the final files. kodwai shows you everything that led there: every prompt, every commit, every test run, every tool call, in the order it happened. Follow it live, or open it after.

claude-code · rate-limiter
rec18:24
session transcript · in order
00:00you ›spec first: per-key sliding window, monotonic clock, no leak
00:31claudescaffolds limiter.py, adds RateLimiter + window store
02:48$ pytest -q
02:55!9 passed, 1 failed
03:20you ›the burst test races. write the failing test, then fix it
05:02claudeadds per-key lock, scopes the critical section
07:41$ pytest -k concurrency
07:49✓4 passed
09:18claudeedits limiter.py: evict stale keys on read
11:36◇commit: handle burst window + per-key lock
18:24you ›
signals
time18:24
cost$0.42
tools7
tests4 / 5
team review
MayaNice, wrote the failing test first before touching the lock.
RafCaught the race in one steer. Strong recovery.
//reviewers can comment on any moment
//04transparent scoring

Score the work, not the memory.

You define the rubric for the role. Every finished session is scored against it, with the reasoning shown for each dimension, and your team scores it too.

Jamie BrooksSenior Backend Engineer
// Sliding-window rate limiter
8.4
/ 10vs your rubric
Problem decomposition30%8.5
Direction and verification30%9.0
Code quality25%8.0
Recovery under failure15%7.5
AI review

A methodical session. Scoped the ticket before touching code, drove the agent with tight, verifiable steps, and recovered cleanly when the first approach to the limiter leaked memory.

strengths
  • Wrote a failing test before each change
  • Caught the agent's off-by-one in the window math
  • Kept commits small and clearly labelled
weaknesses
  • Left clock-skew as an uncovered edge case
  • Did not write down the final trade-off

Compare the role.

rubric heatmap
candidate
Decompose
Direction
Quality
Recovery
overall
Jamie Brooks@jamie
8.5
9.0
8.0
7.5
8.4
Sarah Chen@schen
7.0
7.5
9.0
6.5
7.5
Kenji Tanaka@ktanaka
6.0
6.5
7.0
8.5
6.8

scored 0 to 10 the same rubric for everyone strongest and weakest dimensions surface at a glance

//05the proof

Hiring you can actually see.

No sandbox, no take-home, no guessing. You watch the real work, you score it on your terms, and you compare like for like.

the invite1 link
to invite a candidate. No portal for them to set up and no account to provision.
the machine0%
local. They solve on their own machine, in Claude Code.
the visibilitylive
Watch the prompts, commits, tests, and cost as the session happens, not days later.
the grading0
take-homes to grade at midnight. The session arrives scored, with the full replay.
measured, not promisedthe session is the resume
//06nothing to connect

Nothing to connect.

No ATS to wire up, no stack to sync, no candidate accounts to provision. You send a link, they solve where they already work, and a scored session comes back.

  • One link out, one scored session back.
  • They work in Claude Code, nothing new to learn.
  • Invite your team as admin, interviewer, or viewer.
  • Share any session with a link.
no integration required

Send a link. That is the integration.

the tool they already know
Claude Codethe same CLI as the developer track
  • 01
    Send the invite
    One link, scoped to the challenge and time limit you set. No account to provision, no stack to wire up.
  • 02
    They solve locally
    The candidate works on their own machine in Claude Code, the agent they already know. Same CLI as the developer track.
  • 03
    You review the replay
    A scored session lands back: the prompts, commits, test runs, and transcript next to the score. Share one link with your team.
//07questions

Frequently asked.

Everything worth knowing before you run your first interview. Still curious, the answer is one message away.

How does a candidate take a Kodwai interview?// candidate

You add them by name and email, and they get a link with one command to run. They solve the problem on their own machine in Claude Code, the agent they already know. There is no browser sandbox to fight and nothing to install beyond the CLI.

What can I see while they work?// observe

The full session, live: every prompt, the commits, the test runs, the tools used, the time elapsed, and the API cost. You can follow along in real time or open the session afterward and read it in the order it happened.

How are interviews scored?// scoring

You define the rubric for the role, the dimensions, their weights, and what each one means. Every finished session is scored against that rubric with a written justification, strengths, and weaknesses per dimension. Your team can add manual scores and comments next to the AI score.

Can I use my own problem?// problem

Yes. Author the ticket so it looks like your real codebase, then set the time limit, the budget, and the tools the candidate may use. Reuse it for every candidate so the comparison stays fair.

Do you connect to my ATS?// team

Not yet, and we will not pretend otherwise. There is nothing to wire up: you send a link and a scored session comes back. Share that session with your team with a link, and invite teammates as admins, interviewers, or viewers.

Is the hiring track free?// pricing

Solving public challenges is free for developers. The hiring track is how we keep Kodwai sustainable. Reach out and we will get your team set up.

Still deciding?talk to us→
//08begin

Hire for how they drive AI.

The take-home is dead. Watch the real work, score it on your own rubric, and decide as a team, all from one link.