evals
Run evals on Jetty
Method

A pelican learns to ride

This is a worked example of hill-climbing a runbook on Jetty. We took Simon Willison's pelican-on-a-bicycle prompt, wrote it up as a runbook with a self-critique loop, and ran it across 14 agent and model pairs. Every run is a trajectory you can open.

What you're looking at

Jetty runs a runbook, a markdown file with instructions, constraints and a scoring rubric, inside a pinned sandbox. A coding agent (claude-code, opencode, hermes, gemini-cli) follows it and every step is captured as a trajectory. To improve the result you edit the runbook and run it again.

v1 (May 2026) asked what moves the score more: iterating the runbook or iterating the model. The answer was about the same. Editing the runbook by hand gained six self-score points; swapping the model gained five. v2 (September 2026) re-runs the climb with five current models and adds an external judge, because v1 also showed that self-scores can't be compared across models.

The task

Hand-write a pure-XML SVG of a pelican riding a bicycle. No <image> tags, under 50 KB, viewBox around 800×600. Both subjects must be unmistakable, and the pelican has to be riding the bike rather than floating next to it.

The rubric

Four axes, 0–10 each: Pelican (would a stranger say "pelican"?), Bicycle (two spoked wheels, a real frame), Composition (is it actually riding?) and Polish (line, color, balance). The agent renders its SVG, looks at the PNG, scores it, and redraws, three times per run. It keeps the best round.

The hill climb

Each hill-climb round is one Jetty run. Between rounds, the orchestrator (a small Python script) embeds the previous round's best SVG into the runbook as the new starting point and rewrites the description to target the lowest-scoring axis. v1 ran 10 rounds per agent (3 for the Fusion sweeps); v2 ran 5. The one exception is the reference lineage: Jon read each claude-code + Sonnet trajectory and edited the runbook by hand, eleven times.

The judge (new in v2)

Every round's final drawing, v1 and v2 alike, is re-scored by one fixed vision judge: Claude Opus 5.5 through OpenRouter, temperature 0, three samples averaged, same rubric. See Judge vs self-score for the prompt and the numbers.

Why this prompt

Simon Willison has run the pelican prompt against nearly every major model release since 2024. He picked it on purpose:

They shouldn't be able to draw anything at all. But they can generate code… and SVG is code.

Simon Willison, June 2025

Most importantly: pelicans can't ride bicycles. They're the wrong shape!

Simon Willison

Everyone needs their own benchmark.

Simon Willison

Drawing a pelican on a bicycle in SVG makes a text model reason in code about shape, anatomy and spatial composition all at once. Pelicans can't ride bicycles, so the result can't be copied from training data. And because SVG allows comments, the models often narrate what they're drawing as they go.

Simon's pelican posts ↗

The v1 hand-curated climb, in three steps

Three picks from Jon's eleven runbook edits (claude-code + Sonnet 4.6). The judge scores were added in v2, after the fact.

Big leap · v2 → v3 · self 32 → 34 · judge 25.2 → 24.5
Embed the baseline
v2
v2 · before
v3
v3 · after

The edit: Use the v2 SVG verbatim as the round-1 input. Don't start from scratch.

What happened: Round 1 jumped from 28 to 33, five points just from skipping the cold start. The single highest-leverage edit in the sequence.

Lesson: When a runbook can carry a working artifact as a seed, it should.

Peak · v4 → v5 · self 36 → 37 · judge 24.5 → 26.2
Coordinate-precise asks
v4
v4 · before
v5
v5 · after

The edit: Close the 13 px right-wing-to-grip gap. Drop the left-wing tips to y ≈ 230.

What happened: Composition hit 10/10 for the first time. The pelican grips the bars. Sonnet's peak self-score, never beaten later.

Lesson: At the top of the curve, the asks become measurements.

Cautionary · v5 → v6 · self 37 → 36 · judge 26.2 → 26.5
Scope creep cost a point
v5
v5 · before
v6
v6 · after

The edit: Add motion lines, wind tufts and a sun. Extend the right-wing covert lines to the wrist.

What happened: A richer scene, but composition slipped back to 9. The new elements added clutter the rubric couldn't reward.

Lesson: More isn't better when the rubric doesn't measure scene density.

What we changed for v2

  • Current models, one provider path. Everything routes through OpenRouter with a single collection secret, OPENROUTER_API_KEY, so anyone can rerun it with one key.
  • Five rounds, not ten. Most v1 lineages had plateaued by round five.
  • An external judge over every drawing from both cohorts, so the leaderboard ranks drawings rather than egos.
  • Same runbook. The v2 runbook is v1's baseline with the model defaults updated. The orchestrator logic is unchanged apart from the collection and round count.

Want to try it? Run it yourself.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.