Prepare your repository for the remaining HADR activities. You built the report and incident stores, then built the agent loop. Now add automated quality checks, merge open changes, and record a baseline run.

Do all work on this page in your own repository. Agree on the quality tools with your team. Ask teammates to review the changes that you merge.

Activity: Quality checks

Add code quality checks that run after repeated AI edits. Run them from a pre-commit hook. These checks prevent a gradual reduction in code quality.

Use these standard Python tools:

  1. ruff finds unused, inconsistent, or incorrect code.
  2. pyright or mypy finds type errors.
  3. pytest runs tests.

Add these three (and behave) to your repo and configure them as a pre-commit hook.

Discuss other useful checks with your team. Agree on a shared set of checks. Each team member must add these checks to their own repository.

Some industry standards
  • Test coverage thresholds (pytest --cov --cov-fail-under)
  • Cyclomatic complexity limits (ruff rule C901)
  • Dead-code detection (vulture)
  • Dependency vulnerability audits (pip-audit)
  • Enforced formatting (ruff format --check)

The behave gate

The scenarios in starter/features/ were written before the code. Some scenarios will fail until you implement their features.

  1. Add uv run behave as a pre-commit hook and run it. Each failed scenario identifies work that is not complete. You can bypass a local hook with git commit --no-verify, so the hook is not a quality gate.

  2. Use GitHub CI as the quality gate. Run it for each pull request into master. Initially, run only the scenarios for completed features:

    - run: uv run behave --tags="@warmup"   # gate the store warm-up only
    
  3. Add tags to the gate when you complete each stage. After Stage 1 passes, use --tags="@warmup or @stage1". Add @stage2 after you write the intake scenarios. Add @stage3 after you complete reconciliation. Keep completed stages in the gate to prevent later changes from breaking them.

  4. Add a rule to CLAUDE.md: check the CI tag expression when a feature file changes or a stage is completed.

When you’re done

Post your list of pre-commit hooks to the Padlet. Review the lists from other teams and add useful checks to your list.

Activity: Merging

  1. Merge all open changes into your own master. This includes the phase 1 and phase 2 loop branches and the pull request from the evaluation driver. A teammate must review each pull request before you merge it. Share the reviews across your team.

  2. Use a clean clone or worktree. Run uv run behave --tags=@warmup for the stores. Run uv run behave features/agent_loop.feature for the loop. Then run the skeleton episode with uv run hadr-runner stage1-basic@londone --no-llm. Start the engine first with uv run https://dl.hadr.ocelliq.com/hadr-engine.py serve --port 8000. The skeleton saves nobody. Record this result as the baseline.

  3. Run the live loop on some examples (uv run hadr-agent "what's the secret message?") and record the rounds and token total it prints.

  4. Check the prompt cache on the live run. agent_loop sends each assistant turn back without changes. Thus, the system prompt, tool definitions, and prior turns stay the same across rounds. The kit adds the provider’s prompt_tokens_details.cached_tokens value to result.usage.cache_read. Add cache_read to the final output in main(). Confirm that it is not zero for a run with multiple rounds.

When you’re done

  1. Post your run on the Padlet. Include the rounds, total tokens, and cache_read from step 3. A cache_read of zero can mean that code changes the assistant turn. Compare your code with code from a teammate who has a nonzero value.
  2. Complete the integration. Merge all required pull requests. Confirm that CI passes on master. Install the pre-commit hook in the clone that you will use for the remaining activities.

Leave with

  • A reproducible setup command and test command.
  • Automated checks that maintain code quality.
  • A merged agent loop on your own master: green LOOP-* and store scenarios, plus a live hadr-agent run.
  • The group intake fixtures merged into your repository, ready for Stage 2: intake.