Play before you build. You will direct Claude through two complete episodes while it calls the simulator tools. This gives you examples of the decisions, state, and failures that the automated agent must handle.
Connect
If the simulator is not running, start it in a separate terminal and leave it open:
uv run https://dl.hadr.ocelliq.com/hadr-engine.py serve --port 8000
Open http://127.0.0.1:8000/ to see the public display.
Claude reads MCP configuration when it starts. Open a new Claude Code session from the starter’s mcp-play/ folder:
cd mcp-play
claude
The folder contains three small configuration files:
.mcp.jsonconnects to the simulator athttp://localhost:8000/mcp..claude/settings.jsonselects the model and allows the HADR tools.CLAUDE.mdgives Claude the basic play instructions.
Read these files, then run /mcp. Confirm that the hadr server is connected and that its tools are available. The most recent MCP session controls the simulator, so a new connection takes control from an older client.
The loop
Play is the same lockstep loop the automated runner uses:
list_scenariosin the lobby. Scenario IDs are<template>@<map>pairs, so a listing of five templates over three maps is fifteen entries.start_episodewith a scenario ID (optional seed); it returns the frozen tick-0 observation.- Read contacts and truck vision, use
building_to_coordsandtravel_timeto plan, thendispatchtrucks. next_tick(expected_tick = observation.tick)to advance exactly one tick, and repeat until the observation isterminal.
Only the outer runner (you) calls next_tick. The full tool reference is in the API guide.
Play two episodes
Direct Claude; do not ask it to run unattended yet.
Episode 1: learn the controls
- Ask Claude to list the scenarios and start a stage-1 episode.
- Use every important tool at least once: inspect contacts and trucks, query the map, estimate travel time, dispatch a truck, and advance the tick.
- Before each
next_tick, predict what the next public state will show. - Keep a scratch list of suspected incidents. The simulator does not own this list; it represents your beliefs.
- Continue until the episode is terminal. Record casualties, buildings lost, and one mistake.
Episode 2: improve the procedure
- Restart the same scenario and seed so that the evidence is comparable.
- Give Claude clearer instructions based on the first failure.
- Ask it to state the evidence for each dispatch and what would cause it to redirect a truck.
- Finish the episode and compare the result with Episode 1.
As you play, answer these questions:
- Which state is public, which is an agent belief, and which is hidden truth?
- Why does a dispatch include an
incident_idif the simulator does not validate it? - Why must only the outer runner call
next_tick()? - Which repeated steps should become software, and which still require judgement?
Save the evidence
Keep the transcript and scratch incident list. Write down:
- the per-tick checklist that worked;
- one instruction that improved the second run;
- one failure that the instructions did not prevent; and
- the stopping condition.
Do not write a skill yet. After you build the stores, you will use these two episodes to turn the procedure into a skill.