{
  "init_params": {
    "agent": "claude-code",
    "model": "anthropic/claude-sonnet-4.6",
    "snapshot": "openra-arena",
    "instruction": "---\nversion: \"1.1.0\"\nevaluation: programmatic\nagent: claude-code\nmodel: anthropic/claude-sonnet-4.6\nmodel_provider: openrouter\nsnapshot: openra-arena\nprimary_outputs:\n  - match.mp4\n  - summary.md\n  - result.json\nsecrets:\n  OPENROUTER_API_KEY:\n    env: OPENROUTER_API_KEY\n    description: \"OpenRouter key used by the two competing models (and by this runbook's agent)\"\n    required: true\n---\n\n# OpenRA Arena: one LLM-vs-LLM Red Alert game \u2014 Agent Runbook\n\n## Objective\n\nPlay **one full 1v1 game of Command & Conquer: Red Alert (OpenRA)** between two decision models and deliver the\nreplay, a 1080p recording, the full decision log, and scored metrics to `{{results_dir}}`.\n\nEach player is an LLM (any OpenRouter model) or OpenRA's built-in AI. The game engine blocks every 8 seconds of game\ntime, sends both players their own fog-of-war view **at the same tick**, and resumes only when both have answered, so\nmodel latency never costs game time. A game ends when one side loses all its buildings, or at `{{max_ticks}}` ticks\n(25 ticks = 1 game-second; 30000 = 20 game-minutes), where the side with the lower score (assets value + value of\nenemy stuff destroyed) surrenders.\n\nThe `openra-arena` snapshot ships everything prebuilt: the patched OpenRA engine at `/opt/openra`, Red Alert game\ndata at `/opt/openra-content`, the arena orchestrator at `/opt/decision-ra`, Xvfb + Mesa (software OpenGL) and\nffmpeg. The single entry point is `run-match`. You do not need to build or install anything.\n\n---\n\n## REQUIRED OUTPUT FILES (MANDATORY)\n\n**You MUST write all of the following to `{{results_dir}}`. The task is NOT complete until every file exists and is\nnon-empty.** `run-match` writes all of them itself; your job is to launch it, supervise it, and verify.\n\n| File | Description |\n|------|-------------|\n| `{{results_dir}}/result.json` | Winner, end reason, end tick, final per-player stats, LLM cost/latency/failure counts |\n| `{{results_dir}}/metrics.json` | Normalized per-player metrics (value destroyed/lost, peak army, orders, latency, cost) |\n| `{{results_dir}}/decisions.jsonl` | Every decision turn: the briefing each model saw, its reasoning, commands and their results |\n| `{{results_dir}}/replay.orarep` | OpenRA replay (open it in OpenRA to watch the game with full controls) |\n| `{{results_dir}}/match.mp4` | 1920x1080 recording of the whole game at 6x speed with the observer stats table |\n| `{{results_dir}}/thumb.jpg` | Thumbnail from the recording |\n| `{{results_dir}}/runtime_config.json` | Exact configuration: players, models, map, limits, engine commit + patch hash, host |\n| `{{results_dir}}/summary.md` | Human-readable result table |\n| `{{results_dir}}/validation_report.json` | Programmatic checks with `passed` and per-check results |\n\n---\n\n## Parameters\n\n| Parameter | Template Variable | Default | Description |\n|-----------|-------------------|---------|-------------|\n| Results directory | `{{results_dir}}` | `/app/results` | Output directory (persisted by Jetty) |\n| West player | `{{player_a}}` | `jev` | Player in the west slot (see *Player specs*) |\n| East player | `{{player_b}}` | `sonnet5` | Player in the east slot |\n| Game ID | `{{game_id}}` | `game` | Label for this game (used in logs and the result) |\n| Max ticks | `{{max_ticks}}` | `30000` | Game-length cap in ticks (30000 = 20 game-minutes) |\n\n### Player specs\n\n| Spec | Meaning |\n|------|---------|\n| `jev`, `sonnet5`, `gpt56terra`, `gemini38f` | Roster presets (Jev Router, Claude Sonnet 5, GPT-5.6 Terra, Gemini 3.8 Flash) |\n| `openrouter:<model>` or `openrouter:<model>\\|<Display Name>` | Any OpenRouter chat model, e.g. `openrouter:qwen/qwen3.8-max\\|Qwen 3.8 Max` |\n| `cpu:beginner`, `cpu:easy`, `cpu:medium`, `cpu:normal` | OpenRA's built-in skirmish AI, weakest to strongest |\n\n---\n\n## Dependencies\n\n| Dependency | Where | Notes |\n|------------|-------|-------|\n| OpenRA engine + arena bridge | `/opt/openra` | Built from `yxc20089/OpenRA@9b271c1` + `patches/openra-arena.patch` |\n| Red Alert game data | `/opt/openra-content` | Freeware content (OpenRA quick-install bundle) |\n| Orchestrator | `/opt/decision-ra` (`run-match` on PATH) | Python; calls OpenRouter |\n| `OPENROUTER_API_KEY` | collection env var | Required unless both players are `cpu:*` |\n\n---\n\n## Step 1: Environment Setup\n\n```bash\nmkdir -p {{results_dir}}\ntest -x /usr/local/bin/run-match && test -f /opt/openra/bin/OpenRA.dll && test -d /opt/openra-content/ra/v2 \\\n  && echo \"PASS: openra-arena snapshot present\" || echo \"FAIL: not running on the openra-arena snapshot\"\ntest -n \"${OPENROUTER_API_KEY:-}\" && echo \"PASS: OPENROUTER_API_KEY set\" || echo \"WARN: OPENROUTER_API_KEY missing (only cpu:* players will work)\"\nnproc; free -g | head -2\n```\n\nIf the snapshot check fails, stop and write `validation_report.json` with `passed: false` and the reason. Do not\ntry to build OpenRA inside the sandbox.\n\n## Step 2: Launch the Game (as a detached process)\n\nA game takes roughly 20\u201345 minutes of wall-clock time, longer than a single shell call may block. Start it as a detached process with `nohup` (a plain shell command, not an agent background task)\nand write its console log into the results directory:\n\n```bash\ncd {{results_dir}}\nnohup run-match --a \"{{player_a}}\" --b \"{{player_b}}\" --game-id \"{{game_id}}\" \\\n  --max-ticks {{max_ticks}} --out {{results_dir}} > {{results_dir}}/run.log 2>&1 &\necho $! > /tmp/run-match.pid\nsleep 60; tail -5 {{results_dir}}/run.log\n```\n\nWithin a minute `run.log` should show `Starting ...` and then one line per decision turn:\n`[game] tick  1201 | <A>: assets ... | <B>: assets ... | 6.2s`.\n\n## Step 3: Supervise Until the Game Ends (foreground only)\n\n> **CRITICAL:** the sandbox \u2014 and the game with it \u2014 is destroyed the moment your session ends. You MUST stay in this\n> step until `run-match` has exited. Do NOT use background tasks, `run_in_background`, `tail -f`, cron or any other\n> asynchronous watcher, and do NOT end your turn \"to check back later\": there is no later. Only run the blocking\n> command below, in the foreground, again and again.\n\nRun this exact command. It blocks for at most 9 minutes and prints the latest progress line:\n\n```bash\nPID=$(cat /tmp/run-match.pid)\nfor i in $(seq 1 54); do\n  if ! kill -0 \"$PID\" 2>/dev/null; then echo \"RUN-MATCH EXITED\"; break; fi\n  sleep 10\ndone\ntail -2 {{results_dir}}/run.log\n```\n\n- If the output does not contain `RUN-MATCH EXITED`, run the same command again immediately. Expect 4\u20138 repetitions:\n  a 20-minute game takes about 30\u201350 minutes of wall-clock time, then the video renders for a few minutes.\n- Do not kill a game that is still logging turns (ticks advance at about 600\u20131500 per minute).\n- Only when you have seen `RUN-MATCH EXITED` continue to Step 4.\n\n## Step 4: Handle Failures (max 1 retry)\n\nRead `{{results_dir}}/validation_report.json`.\n\n- `has_winner: false` (engine crashed, or the orchestrator lost the connection): check `engine.log` and `run.log`,\n  then relaunch Step 2 **once** with `--game-id \"{{game_id}}-retry\"`.\n- `video_rendered` false but the game has a winner: rerun only the render:\n  `cd /opt/decision-ra && DISPLAY=:99 .venv/bin/python -m arena.render /tmp/arena/{{game_id}}` then copy\n  `match.mp4` and `thumb.jpg` into `{{results_dir}}`.\n- Never replay a game that already has a winner just to get a different result.\n\n## Step 5: Report\n\n`run-match` already wrote `summary.md`. Append a short **Notes** section to it: anything unusual in `run.log`\n(LLM errors, retries, truncated video) and one sentence on how the game was won, taken from the last few\n`thoughts` entries in `decisions.jsonl` for each player.\n\n## Step 6: Final Checklist (MANDATORY \u2014 do not skip)\n\n### Verification Script\n\n```bash\necho \"=== FINAL OUTPUT VERIFICATION ===\"\nR=\"{{results_dir}}\"\nfor f in result.json metrics.json decisions.jsonl replay.orarep match.mp4 thumb.jpg runtime_config.json summary.md validation_report.json; do\n  if [ -s \"$R/$f\" ]; then echo \"PASS: $f ($(wc -c < \"$R/$f\") bytes)\"; else echo \"FAIL: $f missing or empty\"; fi\ndone\npython3 -c \"import json; r=json.load(open('$R/result.json')); assert r.get('winner'), 'no winner'; print('PASS: winner =', r['winner'], '|', r['end_reason'])\" || echo \"FAIL: result has no winner\"\npython3 -c \"import json; v=json.load(open('$R/validation_report.json')); print('PASS' if v['passed'] else 'FAIL', 'validation', v['checks'])\"\nffprobe -v error -show_entries format=duration -of csv=p=0 \"$R/match.mp4\" && echo \"PASS: video decodes\"\necho \"=== VERIFICATION COMPLETE ===\"\n```\n\n### Checklist\n\n- [ ] The game ran on the `openra-arena` snapshot with the requested players\n- [ ] `result.json` names a winner\n- [ ] `match.mp4` decodes and `validation_report.json` reports `video_complete: true`\n- [ ] `decisions.jsonl` has turns for every LLM player\n- [ ] `summary.md` includes the Notes section\n- [ ] Verification script printed PASS for every line\n\n**If ANY item fails, go back to Step 4. Do NOT finish until all items pass or the single retry is used up; if it is,\nleave `validation_report.json` as written and explain in `summary.md`.**\n\n---\n\n## Changelog\n\n| Version | Change |\n|---------|--------|\n| 1.1.0 | Supervision must be foreground-only: agents that moved polling into background tasks ended their session and lost the game. |\n| 1.0.0 | Initial version: one game per run, 20-minute cap, lockstep decisions every 8 game-seconds. |\n\n## Tips\n\n- Cost is dominated by the two competing models (roughly $0.05\u2013$3 per game per model), not by this runbook's agent.\n- `decisions.jsonl` rows contain the full text briefing each model saw, useful for debugging strange play.\n- To watch the replay with full controls, install OpenRA and open `replay.orarep` (engine build must match; see\n  `runtime_config.json`).\n",
    "vars": {
      "results_dir": "/app/results",
      "player_a": "jev",
      "player_b": "sonnet5",
      "game_id": "game",
      "max_ticks": "30000"
    },
    "file_paths": [],
    "model_provider": "openrouter"
  },
  "steps": [
    "run"
  ],
  "step_configs": {
    "run": {
      "activity": "runbook",
      "agent_path": "init_params.agent",
      "model_path": "init_params.model",
      "snapshot_path": "init_params.snapshot",
      "instruction_path": "init_params.instruction",
      "template_variables_path": "init_params.vars",
      "files_path": "init_params.file_paths",
      "cpus": 4,
      "memory": "8G",
      "timeout_sec": 7200,
      "langfuse_tracing": true,
      "model_provider_path": "init_params.model_provider"
    }
  }
}