evals
Run evals on Jetty
Round robin · rr-sonnet5-gpt56terra-g2

GPT-5.6 Terra vs Claude Sonnet 5

Claude Sonnet 5 destroyed the enemy base after 16.8 game-minutes · 22.4 min wall clock

LOSS ukraine · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$29,700Value lost$48,800 Units killed / lost41 / 87Buildings killed / lost4 / 16 Peak army$7,000Orders issued162 Decision turns (failed)126 (0)Mean latency4.2sModel cost$1.063
WIN germany · east
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$48,800Value lost$29,700 Units killed / lost87 / 41Buildings killed / lost16 / 4 Peak army$8,900Orders issued385 Decision turns (failed)126 (0)Mean latency4.6sModel cost$1.261

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.