---
title: "Run it yourself | OpenRA arena | Evals by Jetty"
url: https://evaljetty.com/openra/run.html
description: "The runbook, sandbox snapshot and API calls to run the OpenRA LLM arena on your own Jetty account."
updated: 2026-09-28
publisher: Jetty (https://jetty.io)
---

Reproduce

# Run the OpenRA arena on your own Jetty

Everything on this site came from one runbook, executed once per game on [Jetty](https://jetty.io/?utm_source=evaljetty&utm_medium=referral&utm_campaign=openra).
The runbook drives a prebuilt sandbox snapshot with the engine, game data, virtual display and orchestrator already installed, so a game starts in seconds.

[Download RUNBOOK.md](https://evaljetty.com/openra/downloads/RUNBOOK.md)
[Task workflow JSON](https://evaljetty.com/openra/downloads/task-workflow.json)
[Snapshot Dockerfile](https://evaljetty.com/openra/downloads/Dockerfile)
[Source (orchestrator + engine patch)](https://evaljetty.com/openra/downloads/decision-ra-src.tar.gz)

## 1. What runs

|  |  |
| --- | --- |
| Runbook | `openra-arena-match` v1.0.0 · agent `claude-code` · `anthropic/claude-sonnet-4.6` via OpenRouter (supervises the game) |
| Sandbox snapshot | `openra-arena` · 4 vCPU, 8 GB · Debian bookworm, .NET 8, OpenRA (OpenRA-RL fork @ `9b271c1` + arena patch), Xvfb + Mesa, ffmpeg, Python 3.11 |
| Per game | 20-minute cap · decisions every 8 game-seconds · map Singles · ≈20–45 min wall clock |
| Outputs | `result.json`, `metrics.json`, `decisions.jsonl`, `replay.orarep`, `match.mp4`, `runtime_config.json`, `summary.md`, `validation_report.json` |
| Secrets | `OPENROUTER_API_KEY` as a collection environment variable |

## 2. Launch games with the API

Set `JETTY_TOKEN` to an API key for your collection and `COLLECTION` to its name. The `openra-arena` snapshot is available to Jetty runbooks.

```
# 1. Create the task in your collection (once)
curl -s -X POST https://flows-api.jetty.io/api/v1/tasks/$COLLECTION \
  -H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" \
  -d "$(jq -n --rawfile rb RUNBOOK.md --slurpfile wf task-workflow.json \
        '{name: "openra-arena-match", workflow: ($wf[0] | .init_params.instruction = $rb)}')"

# 2. Your collection needs OPENROUTER_API_KEY (Settings → Environment, or:)
curl -s -X PATCH https://flows-api.jetty.io/api/v1/collections/$COLLECTION/environment \
  -H "Authorization: Bearer $JETTY_TOKEN" -H "Content-Type: application/json" \
  --data-binary '{"environment_variables": {"OPENROUTER_API_KEY": "sk-or-..."}}'

# 3. Play a game: any OpenRouter model vs another, or vs OpenRA's AI
curl -s -X POST https://flows-api.jetty.io/api/v1/run/$COLLECTION/openra-arena-match \
  -H "Authorization: Bearer $JETTY_TOKEN" \
  -F 'init_params={"vars": {"player_a": "openrouter:anthropic/claude-sonnet-5|Claude Sonnet 5",
                              "player_b": "cpu:normal", "game_id": "my-first-game", "max_ticks": "30000"}}'
```

## 3. Players

Use a preset ID, `openrouter:<model>|<Display name>` for any OpenRouter chat model, or `cpu:<level>` for OpenRA's AI.

| Spec | Name | Model |
| --- | --- | --- |
| jev | Jev Router | typesafe/jev-router |
| sonnet5 | Claude Sonnet 5 | anthropic/claude-sonnet-5 |
| gpt56terra | GPT-5.6 Terra | openai/gpt-5.6-terra |
| cpu-beginner | OpenRA AI (Beginner) | openra/beginner-ai |
| cpu-easy | OpenRA AI (Easy) | openra/easy-ai |
| cpu-medium | OpenRA AI (Medium) | openra/medium-ai |
| cpu-normal | OpenRA AI (Normal) | openra/normal-ai |

## 4. The runbook

```
---
version: "1.1.0"
evaluation: programmatic
agent: claude-code
model: anthropic/claude-sonnet-4.6
model_provider: openrouter
snapshot: openra-arena
primary_outputs:
  - match.mp4
  - summary.md
  - result.json
secrets:
  OPENROUTER_API_KEY:
    env: OPENROUTER_API_KEY
    description: "OpenRouter key used by the two competing models (and by this runbook's agent)"
    required: true
---

# OpenRA Arena: one LLM-vs-LLM Red Alert game — Agent Runbook

## Objective

Play **one full 1v1 game of Command & Conquer: Red Alert (OpenRA)** between two decision models and deliver the
replay, a 1080p recording, the full decision log, and scored metrics to `{{results_dir}}`.

Each player is an LLM (any OpenRouter model) or OpenRA's built-in AI. The game engine blocks every 8 seconds of game
time, sends both players their own fog-of-war view **at the same tick**, and resumes only when both have answered, so
model latency never costs game time. A game ends when one side loses all its buildings, or at `{{max_ticks}}` ticks
(25 ticks = 1 game-second; 30000 = 20 game-minutes), where the side with the lower score (assets value + value of
enemy stuff destroyed) surrenders.

The `openra-arena` snapshot ships everything prebuilt: the patched OpenRA engine at `/opt/openra`, Red Alert game
data at `/opt/openra-content`, the arena orchestrator at `/opt/decision-ra`, Xvfb + Mesa (software OpenGL) and
ffmpeg. The single entry point is `run-match`. You do not need to build or install anything.

---

## REQUIRED OUTPUT FILES (MANDATORY)

**You MUST write all of the following to `{{results_dir}}`. The task is NOT complete until every file exists and is
non-empty.** `run-match` writes all of them itself; your job is to launch it, supervise it, and verify.

| File | Description |
|------|-------------|
| `{{results_dir}}/result.json` | Winner, end reason, end tick, final per-player stats, LLM cost/latency/failure counts |
| `{{results_dir}}/metrics.json` | Normalized per-player metrics (value destroyed/lost, peak army, orders, latency, cost) |
| `{{results_dir}}/decisions.jsonl` | Every decision turn: the briefing each model saw, its reasoning, commands and their results |
| `{{results_dir}}/replay.orarep` | OpenRA replay (open it in OpenRA to watch the game with full controls) |
| `{{results_dir}}/match.mp4` | 1920x1080 recording of the whole game at 6x speed with the observer stats table |
| `{{results_dir}}/thumb.jpg` | Thumbnail from the recording |
| `{{results_dir}}/runtime_config.json` | Exact configuration: players, models, map, limits, engine commit + patch hash, host |
| `{{results_dir}}/summary.md` | Human-readable result table |
| `{{results_dir}}/validation_report.json` | Programmatic checks with `passed` and per-check results |

---

## Parameters

| Parameter | Template Variable | Default | Description |
|-----------|-------------------|---------|-------------|
| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |
| West player | `{{player_a}}` | `jev` | Player in the west slot (see *Player specs*) |
| East player | `{{player_b}}` | `sonnet5` | Player in the east slot |
| Game ID | `{{game_id}}` | `game` | Label for this game (used in logs and the result) |
| Max ticks | `{{max_ticks}}` | `30000` | Game-length cap in ticks (30000 = 20 game-minutes) |

### Player specs

| Spec | Meaning |
|------|---------|
| `jev`, `sonnet5`, `gpt56terra`, `gemini38f` | Roster presets (Jev Router, Claude Sonnet 5, GPT-5.6 Terra, Gemini 3.8 Flash) |
| `openrouter:<model>` or `openrouter:<model>\|<Display Name>` | Any OpenRouter chat model, e.g. `openrouter:qwen/qwen3.8-max\|Qwen 3.8 Max` |
| `cpu:beginner`, `cpu:easy`, `cpu:medium`, `cpu:normal` | OpenRA's built-in skirmish AI, weakest to strongest |

---

## Dependencies

| Dependency | Where | Notes |
|------------|-------|-------|
| OpenRA engine + arena bridge | `/opt/openra` | Built from `yxc20089/OpenRA@9b271c1` + `patches/openra-arena.patch` |
| Red Alert game data | `/opt/openra-content` | Freeware content (OpenRA quick-install bundle) |
| Orchestrator | `/opt/decision-ra` (`run-match` on PATH) | Python; calls OpenRouter |
| `OPENROUTER_API_KEY` | collection env var | Required unless both players are `cpu:*` |

---

## Step 1: Environment Setup

```bash
mkdir -p {{results_dir}}
test -x /usr/local/bin/run-match && test -f /opt/openra/bin/OpenRA.dll && test -d /opt/openra-content/ra/v2 \
  && echo "PASS: openra-arena snapshot present" || echo "FAIL: not running on the openra-arena snapshot"
test -n "${OPENROUTER_API_KEY:-}" && echo "PASS: OPENROUTER_API_KEY set" || echo "WARN: OPENROUTER_API_KEY missing (only cpu:* players will work)"
nproc; free -g | head -2
```

If the snapshot check fails, stop and write `validation_report.json` with `passed: false` and the reason. Do not
try to build OpenRA inside the sandbox.

## Step 2: Launch the Game (as a detached process)

A game takes roughly 20–45 minutes of wall-clock time, longer than a single shell call may block. Start it as a detached process with `nohup` (a plain shell command, not an agent background task)
and write its console log into the results directory:

```bash
cd {{results_dir}}
nohup run-match --a "{{player_a}}" --b "{{player_b}}" --game-id "{{game_id}}" \
  --max-ticks {{max_ticks}} --out {{results_dir}} > {{results_dir}}/run.log 2>&1 &
echo $! > /tmp/run-match.pid
sleep 60; tail -5 {{results_dir}}/run.log
```

Within a minute `run.log` should show `Starting ...` and then one line per decision turn:
`[game] tick  1201 | <A>: assets ... | <B>: assets ... | 6.2s`.

## Step 3: Supervise Until the Game Ends (foreground only)

> **CRITICAL:** the sandbox — and the game with it — is destroyed the moment your session ends. You MUST stay in this
> step until `run-match` has exited. Do NOT use background tasks, `run_in_background`, `tail -f`, cron or any other
> asynchronous watcher, and do NOT end your turn "to check back later": there is no later. Only run the blocking
> command below, in the foreground, again and again.

Run this exact command. It blocks for at most 9 minutes and prints the latest progress line:

```bash
PID=$(cat /tmp/run-match.pid)
for i in $(seq 1 54); do
  if ! kill -0 "$PID" 2>/dev/null; then echo "RUN-MATCH EXITED"; break; fi
  sleep 10
done
tail -2 {{results_dir}}/run.log
```

- If the output does not contain `RUN-MATCH EXITED`, run the same command again immediately. Expect 4–8 repetitions:
  a 20-minute game takes about 30–50 minutes of wall-clock time, then the video renders for a few minutes.
- Do not kill a game that is still logging turns (ticks advance at about 600–1500 per minute).
- Only when you have seen `RUN-MATCH EXITED` continue to Step 4.

## Step 4: Handle Failures (max 1 retry)

Read `{{results_dir}}/validation_report.json`.

- `has_winner: false` (engine crashed, or the orchestrator lost the connection): check `engine.log` and `run.log`,
  then relaunch Step 2 **once** with `--game-id "{{game_id}}-retry"`.
- `video_rendered` false but the game has a winner: rerun only the render:
  `cd /opt/decision-ra && DISPLAY=:99 .venv/bin/python -m arena.render /tmp/arena/{{game_id}}` then copy
  `match.mp4` and `thumb.jpg` into `{{results_dir}}`.
- Never replay a game that already has a winner just to get a different result.

## Step 5: Report

`run-match` already wrote `summary.md`. Append a short **Notes** section to it: anything unusual in `run.log`
(LLM errors, retries, truncated video) and one sentence on how the game was won, taken from the last few
`thoughts` entries in `decisions.jsonl` for each player.

## Step 6: Final Checklist (MANDATORY — do not skip)

### Verification Script

```bash
echo "=== FINAL OUTPUT VERIFICATION ==="
R="{{results_dir}}"
for f in result.json metrics.json decisions.jsonl replay.orarep match.mp4 thumb.jpg runtime_config.json summary.md validation_report.json; do
  if [ -s "$R/$f" ]; then echo "PASS: $f ($(wc -c < "$R/$f") bytes)"; else echo "FAIL: $f missing or empty"; fi
done
python3 -c "import json; r=json.load(open('$R/result.json')); assert r.get('winner'), 'no winner'; print('PASS: winner =', r['winner'], '|', r['end_reason'])" || echo "FAIL: result has no winner"
python3 -c "import json; v=json.load(open('$R/validation_report.json')); print('PASS' if v['passed'] else 'FAIL', 'validation', v['checks'])"
ffprobe -v error -show_entries format=duration -of csv=p=0 "$R/match.mp4" && echo "PASS: video decodes"
echo "=== VERIFICATION COMPLETE ==="
```

### Checklist

- [ ] The game ran on the `openra-arena` snapshot with the requested players
- [ ] `result.json` names a winner
- [ ] `match.mp4` decodes and `validation_report.json` reports `video_complete: true`
- [ ] `decisions.jsonl` has turns for every LLM player
- [ ] `summary.md` includes the Notes section
- [ ] Verification script printed PASS for every line

**If ANY item fails, go back to Step 4. Do NOT finish until all items pass or the single retry is used up; if it is,
leave `validation_report.json` as written and explain in `summary.md`.**

---

## Changelog

| Version | Change |
|---------|--------|
| 1.1.0 | Supervision must be foreground-only: agents that moved polling into background tasks ended their session and lost the game. |
| 1.0.0 | Initial version: one game per run, 20-minute cap, lockstep decisions every 8 game-seconds. |

## Tips

- Cost is dominated by the two competing models (roughly $0.05–$3 per game per model), not by this runbook's agent.
- `decisions.jsonl` rows contain the full text briefing each model saw, useful for debugging strange play.
- To watch the replay with full controls, install OpenRA and open `replay.orarep` (engine build must match; see
  `runtime_config.json`).
```

---
Source: https://evaljetty.com/openra/run.html · Evals by Jetty · Run your own evals: https://jetty.io/?utm_source=evaljetty&utm_medium=llms
