evals
Run evals on Jetty
Champion vs OpenRA AI · ladder-gpt56terra-beginner-g2

OpenRA AI (Beginner) vs GPT-5.6 Terra

OpenRA AI (Beginner) won on score at the time limit (35,200 to 22,250) after 20.0 game-minutes · 27.5 min wall clock

WIN ukraine · west
OpenRA AI (Beginner)
openra/beginner-ai
Value destroyed$26,500Value lost$6,550 Units killed / lost93 / 32Buildings killed / lost3 / 3 Peak army$0Orders issued62 Decision turns (failed)0 (0)Mean latency–Model cost$0.000
LOSS england · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$6,550Value lost$26,500 Units killed / lost32 / 93Buildings killed / lost3 / 3 Peak army$3,200Orders issued137 Decision turns (failed)151 (0)Mean latency3.7sModel cost$1.183

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.