Build an agent that turns imperfect emergency reports into fire-truck dispatches. Start by playing the game through Claude so that the problem is concrete. Then build the agent one layer at a time.

Each participant works in a separate HADR repository. Teammates review pull requests and share the evaluation work.

Understand the task

  1. The HADR game: learn the mission, information boundaries, and system shape.
  2. MCP Play: direct Claude through two complete episodes before you inspect or automate the pipeline.
  3. Starter kit: locate the files, tests, and incomplete components.

Build the foundations

  1. Plan, build, review: implement the report and incident stores from the supplied specification.
  2. Turn play into a skill: extract a reusable procedure from the episodes you already played.
  3. Tokens and costing: inspect context growth, token use, and prompt caching.
  4. Subagents and evals: produce reviewed intake fixtures for later evaluation.
  5. Agentic loop: build and test the task-neutral tool loop.
  6. Quality gates and integration: add automated checks, merge the work, and record a baseline.

Complete and evaluate the pipeline

  1. Stage 1: dispatcher: configure the loop with state, instructions, and typed tools.
  2. Stage 2: intake: convert human text into structured reports and evaluate it.
  3. Stage 3: reconciliation: integrate deterministic matching for duplicate and conflicting reports.
  4. Stage 4: replanning: make one measured change for new evidence.
  5. Model comparisons: compare models against the same evidence.
  6. HADR challenge: run the announced assessment game and publish the terminal result.

The aim is not to add every possible HADR feature. The aim is to understand the complete system and improve it with evidence.