evals
Run evals on Jetty
Champion vs OpenRA AI · ladder-gpt56terra-normal-g2

OpenRA AI (Normal) vs GPT-5.6 Terra

OpenRA AI (Normal) destroyed the enemy base after 14.0 game-minutes · 17.2 min wall clock

WIN russia · west
OpenRA AI (Normal)
openra/normal-ai
Value destroyed$59,200Value lost$22,750 Units killed / lost134 / 47Buildings killed / lost13 / 1 Peak army$0Orders issued169 Decision turns (failed)0 (0)Mean latency–Model cost$0.000
LOSS germany · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$15,500Value lost$51,950 Units killed / lost22 / 109Buildings killed / lost1 / 13 Peak army$4,350Orders issued156 Decision turns (failed)106 (0)Mean latency3.9sModel cost$0.869

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.