evals
Run evals on Jetty
Champion vs OpenRA AI · ladder-gpt56terra-easy-g2

OpenRA AI (Easy) vs GPT-5.6 Terra

OpenRA AI (Easy) destroyed the enemy base after 16.1 game-minutes · 24.3 min wall clock

WIN england · west
OpenRA AI (Easy)
openra/easy-ai
Value destroyed$45,250Value lost$13,450 Units killed / lost101 / 32Buildings killed / lost11 / 1 Peak army$0Orders issued135 Decision turns (failed)0 (0)Mean latency–Model cost$0.000
LOSS england · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$13,450Value lost$45,250 Units killed / lost32 / 101Buildings killed / lost1 / 11 Peak army$3,900Orders issued136 Decision turns (failed)122 (0)Mean latency3.6sModel cost$0.946

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.