3 rounds of dev screening,
automated into 1 challenge.

hunr puts every developer on real work — and shows you who actually understands what they ship.

Evidence-linked scores, benchmarked against a senior engineer.

Same task · same score0evaluated
Engineer A0/14 tests
Code score0/ 100
Identical on paper
Engineer B0/14 tests
Code score0/ 100
1
Challenge

A real task from a real codebase — not a puzzle.

From your JD, hunr generates a runnable challenge in the candidate's actual stack: real repo, hidden tests, a rubric weighted to what you care about.

See it for your team
  • Role- and stack-specific, not a generic LeetCode screen
  • Hidden tests are authoritative; you approve before it goes live
  • Your competency weights drive how candidates are ranked
2
Build — agents allowed

Let them use AI. That's the whole point.

Candidates work in their own repo with any agent. Every push runs against the hidden gateway suite in an isolated sandbox, verdict streamed live. Real workflow, not a locked-down editor.

  • Visible tests give GitHub-native feedback while they work
  • Hidden tests are authoritative and can't be shadowed
  • Reproducible, hard to fake, mirrors the real job
3
Defend — unaided

Then prove you understood it.

A live reasoning round asks why — trade-offs, failure modes, the road not taken. An ownership multiplier discounts the score when answers don't hold up. The gap between a passing artifact and real understanding is the signal.

For candidates
  • Question types: rationale, trade-off, failure mode, novel extension
  • Ownership factor multiplies the technical score; below the floor, it fails
  • One question per turn — time-boxed, asked unaided, scored live
4
Score

A JD in. A ranked shortlist out.

Every candidate returns as a multi-dimensional, evidence-linked fit score — banded Strong / Promising / Borderline / Weak, anchored to a reference expert. The honest signal, not false precision.

  • Per-dimension badges: design, code quality, security, ownership
  • Anchored to a reference expert so a score means something
  • Fast, confident decisions you can defend to your team

Real evaluation has always been a luxury

Take-homes and interviews measure ability — but they're too slow to run on everyone, so they only reach the few the résumé already approved.

2,000 applicants · one role
0
of 2,000 evaluated0% coverage

2,000 never really looked at

Todaydrag to compareWith hunr

Everyone else either bans the agent or babysits it

Use any agent you want. Then prove you understood what it built. The gap between the two is the only thing worth scoring.

The detect-and-block camp

Proctoring, lockdown browsers, AI-detection. An arms race against tools every candidate already uses daily.

What it measures

Can you hide the AI

The watch-you-use-it camp

Interview AI that observes your prompts but can't run code, edit files, or ship anything real.

What it measures

How you talk to a model

hunr

Use any agent to build the thing. Then face a scored gate that proves you actually understood it.

What it measures

Whether you understand the work

Hard to game. Fair to candidates.

Open take-home, unaided defense, scores anchored to real references, and a report every candidate can challenge. Fairness isn't a policy here — it's in the mechanics.

Banded, not false-precise

73 vs. 71 is noise. The band is the honest signal.

GitHub never penalizes

Profile analysis is bonus-only and capped — it corroborates, never condemns.

Full transparency

Every candidate gets the full score and an evidence-linked report.

Adaptive, unaided defense

One question per turn, time-boxed and grounded in your own code.

The best developer rarely has the best résumé. Stop hiring on paper.

Use any agent. Then prove you understood the work. Book a demo, or take a challenge yourself.