---
title: "Case study: building the OpenRA arena on Jetty | Evals by Jetty"
url: https://evaljetty.com/openra/case-study.html
description: "How a real-time strategy LLM benchmark became one runbook, one sandbox snapshot and dozens of parallel Jetty runs, with receipts for every game."
updated: 2026-09-28
publisher: Jetty (https://jetty.io)
---

Case study · built on Jetty

# From a weekend idea to a reproducible benchmark

How the OpenRA arena went from “can an LLM play Red Alert?” to 20 fan-out runs with videos, decision logs and receipts,
with one runbook and one sandbox snapshot on Jetty.

Jetty runs

20

one game per run, in parallel

Sandbox time

5 h

4 vCPU / 8 GB each

Game time played

3.5 h

2,266 model decisions

Model spend

$15

competing models, via OpenRouter

## The problem

A real-time strategy game is a hard, honest test for an agent: partial information, a clock that doesn't wait, an economy to run and an
opponent adapting to you. But an RTS eval is also an infrastructure headache: a C# game engine, native graphics, game data, an orchestrator,
and games that take the better part of an hour each. Running a tournament means dozens of those in parallel, and people only trust the result
if they can see every game and rerun it themselves.

## What we built

1. **An engine bridge.** A bot trait inside OpenRA pauses the simulation every 8 game-seconds, hands both players their
   fog-of-war view at the same tick, and resumes when both models answer. Lockstep keeps it fair and makes every game replayable.
2. **A sandbox snapshot.** The engine, Red Alert data, a virtual display with software OpenGL, ffmpeg and the Python orchestrator
   are baked into one image, registered as the Jetty snapshot `openra-arena`. A run starts playing within seconds instead of spending
   ten minutes compiling.
3. **A runbook.** [One markdown file](https://evaljetty.com/openra/run.html) tells the agent how to launch a game, supervise it, recover from
   failures and verify the outputs: result, metrics, decision log, replay, 1080p video, runtime config and a validation report.
4. **Fan-out.** Each game is one API call with different parameters. The round robin was 12 calls; the champion's ladder against
   OpenRA's built-in AI was 8 more. Jetty ran them side by side and kept every trajectory.
5. **A results site generated from the receipts.** This site is built from the artifacts the runs saved, so every number links back
   to the run that produced it ([see every run](https://evaljetty.com/openra/runs.html)).

## What we learned

- **Replay determinism is a feature you have to test.** Early pilot videos drifted out of sync. The cause: the bridge sent building
  sales as “immediate” orders that bypass the lockstep queue, so the replay applied them at a different tick. Every desync lined up with a sale;
  routing sales through the normal queue fixed it. A validation check (`video_complete`) now catches any recurrence.
- **Tell the agent what “waiting” means.** In the first fan-out, a few runbook agents moved their polling into background tasks and
  ended their session, which tears down the sandbox mid-game. Runbook v1.1.0 makes supervision foreground-only and says why; the lost runs were
  simply relaunched with the same parameters.
- **Bake, don't build.** Moving the engine build into the snapshot cut each run's setup from minutes to seconds and removed a whole
  class of flaky failures.
- **Keep the receipts machine-readable.** Because each run writes `result.json`, `metrics.json` and
  `runtime_config.json`, adding a new chart (like the radar profiles) is a site rebuild, not a re-run.

## Run your own

Swap in your own models, maps or limits and point the same runbook at your collection. If you have an eval that needs a real environment
(a game, a browser, a codebase, a GPU), the pattern is the same: bake the environment, write the runbook, fan out, publish the receipts.

[Start free on Jetty](https://jetty.io/?utm_source=evaljetty&utm_medium=referral&utm_campaign=openra-case-study)
[See the runbook](https://evaljetty.com/openra/run.html)

---
Source: https://evaljetty.com/openra/case-study.html · Evals by Jetty · Run your own evals: https://jetty.io/?utm_source=evaljetty&utm_medium=llms
