Play before you build. You will direct Claude through two complete episodes while it calls the simulator tools. This gives you examples of the decisions, state, and failures that the automated agent must handle.

Connect

If the simulator is not running, start it in a separate terminal and leave it open:

uv run https://dl.hadr.ocelliq.com/hadr-engine.py serve --port 8000

Open http://127.0.0.1:8000/ to see the public display.

Claude reads MCP configuration when it starts. Open a new Claude Code session from the starter’s mcp-play/ folder:

cd mcp-play
claude

The folder contains three small configuration files:

  • .mcp.json connects to the simulator at http://localhost:8000/mcp.
  • .claude/settings.json selects the model and allows the HADR tools.
  • CLAUDE.md gives Claude the basic play instructions.

Read these files, then run /mcp. Confirm that the hadr server is connected and that its tools are available. The most recent MCP session controls the simulator, so a new connection takes control from an older client.

The loop

Play is the same lockstep loop the automated runner uses:

  1. list_scenarios in the lobby. Scenario IDs are <template>@<map> pairs, so a listing of five templates over three maps is fifteen entries.
  2. start_episode with a scenario ID (optional seed); it returns the frozen tick-0 observation.
  3. Read contacts and truck vision, use building_to_coords and travel_time to plan, then dispatch trucks.
  4. next_tick(expected_tick = observation.tick) to advance exactly one tick, and repeat until the observation is terminal.

Only the outer runner (you) calls next_tick. The full tool reference is in the API guide.

Play two episodes

Direct Claude; do not ask it to run unattended yet.

Episode 1: learn the controls

  1. Ask Claude to list the scenarios and start a stage-1 episode.
  2. Use every important tool at least once: inspect contacts and trucks, query the map, estimate travel time, dispatch a truck, and advance the tick.
  3. Before each next_tick, predict what the next public state will show.
  4. Keep a scratch list of suspected incidents. The simulator does not own this list; it represents your beliefs.
  5. Continue until the episode is terminal. Record casualties, buildings lost, and one mistake.

Episode 2: improve the procedure

  1. Restart the same scenario and seed so that the evidence is comparable.
  2. Give Claude clearer instructions based on the first failure.
  3. Ask it to state the evidence for each dispatch and what would cause it to redirect a truck.
  4. Finish the episode and compare the result with Episode 1.

As you play, answer these questions:

  • Which state is public, which is an agent belief, and which is hidden truth?
  • Why does a dispatch include an incident_id if the simulator does not validate it?
  • Why must only the outer runner call next_tick()?
  • Which repeated steps should become software, and which still require judgement?

Save the evidence

Keep the transcript and scratch incident list. Write down:

  • the per-tick checklist that worked;
  • one instruction that improved the second run;
  • one failure that the instructions did not prevent; and
  • the stopping condition.

Do not write a skill yet. After you build the stores, you will use these two episodes to turn the procedure into a skill.