evals
Run evals on Jetty
Pilot cup · cup1-third-g1

Gemini 3.8 Flash vs GPT-5.6 Terra

GPT-5.6 Terra destroyed the enemy base after 4.5 game-minutes · 5.1 min wall clock

Pilot game (run locally, before the Jetty runs) Replay (.orarep) Decision log Runtime config
LOSS england · west
Gemini 3.8 Flash
google/gemini-3.8-flash
Value destroyed$2,200Value lost$10,900 Units killed / lost16 / 16Buildings killed / lost0 / 6 Peak army$1,900Orders issued35 Decision turns (failed)35 (0)Mean latency3.2sModel cost$0.105
WIN france · east
GPT-5.6 Terra
openai/gpt-5.6-terra
Value destroyed$10,900Value lost$2,200 Units killed / lost16 / 16Buildings killed / lost6 / 0 Peak army$4,400Orders issued61 Decision turns (failed)35 (0)Mean latency3.3sModel cost$0.241

Decision timeline

What each model saw as its situation and what it ordered, every 8 game-seconds. Click a turn to jump the video there.

Want this kind of eval for your own agents?

Every run on this site is a Jetty runbook: versioned instructions, a pinned sandbox, and a trajectory you can inspect. Point Jetty at your task and get the same receipts.