Stage 4 changes the evidence after trucks start to move. A responder can correct a headcount. A report can be contradicted. A larger fire can start while both trucks are busy.

A fire truck redirected mid-route to a higher-priority fire as new evidence arrives

The design

Do not add a second planner. The same dispatch loop runs again on each tick.

flowchart LR
    H["next tick"] --> S["new belief state"] --> L["same dispatch loop"] --> A["keep or redirect"] --> H

Replanning needs two things:

  • The state message must show the changed beliefs.
  • The dispatch prompt must permit a useful redirect.

A rule that prohibits all redirects keeps bad early decisions. A rule that redirects too often prevents trucks from reaching fires.

What is provided

The existing system already provides:

  • a loop that runs on every tick;
  • updated incident beliefs from reconciliation;
  • a dispatch tool that can retarget a truck; and
  • the stage4-replan scenario.

What Claude implements

Claude makes one focused improvement. Claude does not rewrite the dispatcher.

Claude changes only these parts:

  1. Claude records a baseline run.
  2. Claude finds one decision that should change after new evidence.
  3. Claude changes _state_message or prompts/dispatch_system.md.
  4. Claude adds one focused scenario for that decision.
  5. Claude tags the scenario @stage4 and updates the CI gate.
  6. Claude compares the result with the baseline.

What you decide and check

You approve the redirect rule before Claude implements it.

The rule must state:

  • which current task a truck can leave;
  • which new incident justifies the redirect;
  • why the new incident is more important; and
  • how the system prevents repeated back-and-forth redirects.

Read one redirect rationale. It must name both sides of the trade. “Going to the highest-priority incident” is not sufficient.

What you do not build

Do not add:

  • a second agent;
  • a new planner;
  • new reconciliation rules;
  • several prompt changes at the same time; or
  • a token limit that exists only in a prompt.

The engine enforces the episode tick limit. Code must enforce any model-spend limit.

Prompt Claude

Paste this prompt into Claude Code:

Improve Stage 4 with one measured replanning change.

Read _state_message, prompts/dispatch_system.md, the incident belief fields, the dispatch tool, the stage4-replan scenario, and the Stage 4 reference notes.

First, run stage4-replan@londone. Record the score, dispatch tokens, and run-log path. Do not edit yet.

Find one point where new evidence should change an earlier truck assignment. Show the state message and dispatch decision before and after the evidence changed.

Propose one small change to _state_message or prompts/dispatch_system.md. The redirect rule must name the task that a truck leaves, the task that it will serve, and why the trade is useful. The rule must also prevent repeated back-and-forth redirects.

Show me the rule and wait for approval.

After approval:

  • Implement only the approved change.
  • Add one focused scenario for the redirect.
  • Run the focused checks.
  • Run the same live scenario again.

Compare casualties, buildings lost, dispatch tokens, and one rationale. If results vary, repeat the scenario before you claim an improvement.

At the end, recommend that we keep or revert the change. Give the evidence for your recommendation.

Check the result

Run:

uv run behave
uv run hadr-runner stage4-replan@londone

Check:

  • New evidence can change a dispatch decision.
  • The rationale names the task that the truck leaves.
  • The truck does not change direction repeatedly.
  • The before-and-after evidence supports the change.

Tune the complete pipeline

After all four stages work, ask Claude to tune one prompt. Paste:

Run the HADR scenarios in stage order. Make a table with the score, tokens, and first important failure for each run.

For each failure, inspect this path:

  1. original contact;
  2. finalized report;
  3. reconciled incident;
  4. dispatch state message;
  5. tool calls; and
  6. final outcome.

Classify each failure as intake, reconciliation, dispatch state, or dispatch prompt. Do not fix an intake error in the dispatch prompt. Do not replace reconciliation code with model judgement.

Choose one high-value prompt change. State a hypothesis. Change only the intake prompt or the dispatch prompt. Run the same evidence again. Compare before and after. Stop if the evidence does not support the change.

Completion criteria

  • One decision changes after the evidence changes.
  • One focused scenario checks the redirect.
  • One prompt or state change has before-and-after evidence.
  • The complete pipeline is ready for model comparison.