evals
Run evals on Jetty
Round robin · rr-sonnet5-gpt56terra-g4

GPT-5.6 Terra vs Claude Sonnet 5

GPT-5.6 Terra destroyed the enemy base after 9.2 game-minutes · 13.7 min wall clock

WIN russia · west
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$20,100Value lost$7,400 Units killed / lost52 / 46Buildings killed / lost11 / 2 Peak army$7,800Orders issued95 Decision turns (failed)69 (0)Mean latency4.0sModel cost$0.553
LOSS ukraine · east
Claude Sonnet 5
anthropic/claude-sonnet-5
Value destroyed$7,400Value lost$20,100 Units killed / lost46 / 52Buildings killed / lost2 / 11 Peak army$3,100Orders issued137 Decision turns (failed)69 (0)Mean latency4.8sModel cost$0.499

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.