live on DevHunt · tool of the week contestcheck it out

Agentic coding challenges: Prove you can drive the agent.

kodwai is a platform where developers solve real coding challenges on their own machine with their own AI coding agent (Claude Code, Cursor, or Codex).

or start from your terminal
  • free for developers
  • runs on your machine
  • claude code · cursor · codex
kodwai on Product Hunt
hard · 110 min · 3 hidden traps
claude code · multi-tenant-rate-limiter-quota-servicerecording
npx @kodwai/cli@latest challenge multi-tenant-rate-limiter-quota-service
--/45
Direction
Direction: not scored yet, out of 45
--/45
Rubric
Rubric: not scored yet, out of 45
--/10
Lift
Lift: not scored yet, out of 10
00total / 100
brilliant: Brilliant redirect
Stopped the agent from evicting keys after 60s, which would have reset a tenant's daily quota.
good: Checked the agent's work
Wrote test_denied_not_recorded before the limiter code changed.
miss: Walked into a trap
Rate and quota are updated in two separate steps, so a denied call still uses up quota.
key moments show up here after scoring
Example run on a real challenge. The scores are made up. The maximum points for each axis are real. Result: Direction 39 of 45, Rubric 37 of 45, Lift 7 of 10, total 83 of 100.
same challenge, two sessions

Same agent.Different driver.

Same challenge, same agent. One developer pasted the spec and walked away. The other planned, tested and pushed back. The tests barely differ. The scores don't.

Hands-off versus hands-on, same challenge
bookshelf-rest-apihands-offhands-on
tests passed11/1214/14
traps caught1/33/3
direction11/4541/45
total4490
session Ahands-offtests 11/12
claude code · bookshelf-rest-apiscored
you ›Here's PROBLEM.md. Build it.
agent ›Done. All endpoints implemented. 11 of 12 tests pass.
npx @kodwai/cli@latest submit
no more prompts
11/45
Direction
Direction: 11 of 45
31/45
Rubric
Rubric: 31 of 45
2/10
Lift
Lift: 2 of 10
44total / 100
dubious: One-shot and pray
No real follow-up after the first prompt.
miss: Walked into a trap
A PATCH with the invalid status "finished" was saved without an error.
Pasted the spec, waited, submitted.
session Bhands-ontests 14/14
claude code · bookshelf-rest-apiscored
you ›Before code: list each endpoint and its validation rules.
agent ›GET/POST /books, GET/PATCH/DELETE /books/:id. Rules below.
you ›PATCH validates like POST. Add a test with a bad status.
agent ›test_patch_rejects_bad_status added. Failing. Fixing.
you ›Filters combine with AND, and q is case-insensitive. Add tests for both.
agent ›2 tests added. 14 of 14 pass.
npx @kodwai/cli@latest submit
41/45
Direction
Direction: 41 of 45
40/45
Rubric
Rubric: 40 of 45
9/10
Lift
Lift: 9 of 10
90total / 100
brilliant: Caught the traps
Asked for a test with an invalid PATCH status before the agent changed the handler.
good: Spec first
Wrote down the validation rules before the first endpoint.
Wrote a spec, added tests, pushed back twice.
how it works

Two commands.One score.

No browser sandbox. Your machine, your editor, the agent you already pay for.

  1. CH.01in your terminal

    Start

    Pick a challenge and run one command. The CLI downloads the problem, starter files and tests, then starts the timer.

  2. CH.02in your editor

    Solve

    Work the way you normally do. Plan, prompt, test and push back when the agent gets it wrong.

    • Claude Code
    • Cursor
    • Codex
  3. CH.03in your terminal

    Submit

    Send your code, commits, test runs and the agent transcript. You get a score out of 100 and the moments that moved it.

the score

Scored on howyou direct the agent.

Passing tests is not enough. Direction weighs the most, because it is the part one lucky prompt can't fake.

DIR
Direction
up to 50 points
--/50
direction
direction: not scored yet, out of 50

How you steer, check and break down the work.

  • Spec Precision
  • Verification Rigor
  • Decomposition
  • Recovery
  • Intent Fidelity
  • Engagement
OUT
Outcome
up to 35 points
--/35
outcome
outcome: not scored yet, out of 35

What shipped: tests and code quality, or the challenge's own rubric.

  • Tests
  • Code Quality
  • Complexity
LIFT
Lift
up to 15 points
--/15
lift
lift: not scored yet, out of 15

The hidden traps you caught, and how far you beat a solo AI.

  • Edge-Case Coverage
  • Lift over AI

Default weights shown. Challenges with their own rubric use 45 / 45 / 10.

Then your session, annotated.

Every scored run lists up to 6 moments that moved your score, chess style, each with its evidence. If you made mistakes, at least 2 of them stay on the list.

  • !! brilliant
  • ✓ good
  • ?! dubious
  • ?? miss
brilliant: Brilliant redirect recovery
"No. 60s resets the daily quota. Evict after the longest window."
miss: Shipped it unread verification_rigor
Accepted a 200-line diff and submitted without running the tests.
challenges

Real tasks.Hidden traps.

Realistic tasks with real tests and a few hidden traps: requirements that are easy to miss if you just paste the spec.

15 challenges · 3 easy · 3 medium · 9 hard

easybackend · 60 minleague points ×1.0

Bookshelf REST API

bookshelf-rest-api

Junior Backend Engineer interview. Build a small REST API from scratch with CRUD, filters, validation, persistence, and a test suite that would catch a regression next sprint.

New here? Start with this one. It is the easy one.

read the full challenge
weekly leaguethis week

A race you canwin this week.

A new league starts every Monday. Your first scored run puts you in a group of up to 30. The top 5 move up a division, the bottom 5 move down.

Points are your best score on each challenge, times its difficulty.easy ×1.0medium ×1.5hard ×2.0
divisions
  1. Bronze
  2. Silver
  3. Gold
  4. Platinum
  5. Diamond
  6. Master
gold division14 players · your rank 12
  1. 1sample-a4 runs341.5
  2. 2sample-b3 runs296
  3. 3sample-c3 runs270
  4. 4sample-d2 runs251.5
  5. 5sample-e3 runs236
  6. 6sample-f2 runs204
  7. 7sample-g2 runs187.5
  8. 8sample-h2 runs150
  9. 9sample-i1 run131
  10. 10sample-j1 run96
  11. 11sample-k1 run88.5
  12. 12you1 run72
  13. 13sample-l1 run60
  14. 14sample-m1 run45
your points so far: bookshelf-rest-api · easy · 72 × 1.0

tiers and profile

Leagues reset.Tiers stick.

Tiers follow your Direction Rating.

Your tier comes from your Direction Rating, an Elo-style number based on how well you direct the agent. Leagues start over every week, but your rating and tier carry over. Grandmaster starts at 1800.

  1. Bronze
    0 to 999
  2. Silver
    1000 to 1149
  3. Gold
    1150 to 1299
  4. Platinum
    1300 to 1449
  5. Diamond
    1450 to 1599
  6. Master
    1600 to 1799
  7. Grandmaster
    1800+

minimum direction rating for each tier · the grandmaster ring uses all three axis colors

public profileapp.kodwai.com/developers/your-username
yu
@your-username
Gold · direction rating 1238
agents
Claude Code
badges
  • First Blood
  • Week Warrior
  • Claude Master
  • Top 10%

Your profile at app.kodwai.com/developers/your-username shows your scores, tier, badges and the agents you use. Every scored run also gets a share card. Send it to the friend who thinks they prompt better than you.

kodwai
90
/ 100 · @your-username
Bookshelf REST API
bookshelf-rest-api
Direction41/45
Rubric40/45
Lift9/10
brilliant: Caught the traps
Asked for a test with an invalid PATCH status before the agent changed the handler.

Every scored run gets a share card with your total, your three axis scores and your best moment.

FAQ

Commonquestions.

Something else?

hakan@kodwai.com
  1. Agentic coding (also called AI-native coding, or vibe coding) is building software by directing an AI agent instead of typing every line. kodwai is a platform where developers solve real coding challenges on their own machine with their own AI coding agent (Claude Code, Cursor, or Codex) and get scored on how well they direct the agent, across three axes: Direction, Outcome, and Lift.

  2. Claude Code, Cursor or Codex. The CLI asks which one. You solve on your own machine, and the CLI only collects from the challenge folder.

  3. Yes for developers. Your first 3 submissions are scored on our key. After that you add your own Anthropic API key in Settings for unlimited runs. It is encrypted and only used to score your work.

  4. Not really. Direction weighs the most, and a session with no spec, no checks and no redirects scores low there however green the tests are. The hidden traps catch the rest.

  5. A new league starts every Monday at 00:00 UTC. Your first scored run puts you in a group of up to 30. Points are your best score on each challenge, times 1, 1.5 or 2 for easy, medium or hard. The top 5 move up a division and the bottom 5 move down.

for hiring teams

Interview the way they really work.

Run a private interview on your own problem and see how the candidate really works, not just the final code.

  • claude code on their own machine
  • scored 0 to 10 on your rubric
  • full transcript of every step
see how interviews work
get started

Skip the puzzles.Build something real.

Pick a challenge, solve it with your agent and see your score. Your first 3 submissions are scored for free.

Start a challenge
or start from your terminal