← about harness-eval

🏆 Demo leaderboard static snapshot

Leaderboard — one PRD, framework is the only variable

Re-weight (client-side, same formula as the CLI)

Per-trial telemetry

Runs across apps

The harness builds the same spec with every framework — and it now runs many specs. Each run is graded against its own frozen PRD and test plan; cross-PRD scores aren't pooled (different specs, different difficulty).

Beyond this snapshot — the full Eval Studio

This page is a read-only leaderboard. The local studio (bun run studio) adds the whole workflow:

Cross-harness

Rank not just frameworks but harnesses — Claude Code, the Codex CLI, and a bare baseline — on one board. Candidates that don't support the chosen harness grey out with the reason.

Configure & launch

Pick target, harness, frameworks, worker/grader models, sandbox provider, trials and concurrency — then launch a real run, a zero-spend dry run, or copy the CLI command.

Live runs

Watch each run stream provisioning → installing → building → grading, with cost-so-far and a cancel that tears down the in-flight sandbox.

Run scorecards

Per-candidate composite + dimensions, a step-by-step adherence comparison, excluded trials with reasons, and frozen provenance hashes.

Trial drill-down

Every PRD test-plan step with pass/partial/fail, weighted credit and cited evidence, plus the blind judge's five criteria with samples.

Boot the artifact

A read-only audit of what the agent built (file tree, cold-start contract, scrubbed blind copy) plus a one-click live demo — boot the built app in a sandbox and get a localhost URL to click through.

Conversation & live stream

Replay the full build transcript with session jumps, an error navigator, and an outline trace — or watch it stream live (redacted) as the build runs.

Pluggable

Five sandbox providers, swappable worker/judge models, multiple harnesses, design-system adherence scoring, and bring-your-own PRD targets.

Run scorecard with dimension bars, trials table, step comparison and excluded trials
Run scorecard — every candidate, side by side, with step-level adherence.
Artifacts panel: read-only audit of the built files plus a one-click live demo with a sandboxed localhost URL
Boot the artifact — audit the build, then run it via a one-click live demo.
Conversation viewer with an error navigator stepping through the run's errors
Conversation & trace — replay the build (or watch it live) and step through errors.
See the full walkthrough → Run it yourself