A model can give different answers when you ask the same question two times. Both answers can seem correct. How can you test the model?

In this section, you will use subagents to make evals for the intake model. An eval is a set of test cases for a large language model (LLM). Each test case contains:

  • An input.
  • An expected output.

Another LLM checks whether the expected output correctly represents the input.

The intake model reads each incoming contact. A contact can be:

  • A call.
  • An alarm.
  • A responder update.

The model converts each contact into a structured report. The dispatch system depends on this report. An intake error can cause errors in the rest of the system.

Run uv run https://dl.hadr.ocelliq.com/hadr-engine.py gen-datasets to get the raw, unlabeled data. The command writes datasets/contacts_unlabeled.jsonl. This file is the input set. The activity in this section creates the expected output set.

Open the file. Examine it before you continue.

Goal: Extract only what the user said

The dispatch software requires structured data. The intake model supplies this data. It extracts:

  • The incident type.
  • The severity.
  • The headcount.
  • The reported evidence.

The model must extract what the source said. It must not decide whether the statement is true.

A caller can name the wrong building or exaggerate the incident. The simulation handles these errors, not the intake model. Each contact already contains these machine-readable values:

  • reported_location
  • occurred_at

Copy these values. Do not derive them again from the contact text.

Each row contains its contact in the contact key. Read row["contact"], not row. For example, a row can contain:

{
  "contact": {
    "contact_id": "unlabeled-con-1",
    "delivered_at": 6,
    "occurred_at": 5,
    "payload": {"text": "Listen, by the harbour market <l139>, 2 people definitely still in there and really bad, there's a fire. There's smoke."},
    "payload_type": "human_text",
    "reported_location": {"building_id": "l139", "kind": "building_id"},
    "source_type": "public_call"
  }
}

The expected output can be:

{
  "incident_type": "fire",
  "severity": "high",
  "headcount": 2,
  "evidence": ["smoke", "flames reported"]
}

You must design the report schema. The schema must define:

  • Its fields.
  • The permitted values for each field.
  • A value for information that the source did not provide.

The activity tests your schema against 300 contacts.

Subagents

A subagent is a separate Claude process. It has its own context window and tool limits. It performs a specified task and returns only the result to the main session.

Subagents have four benefits:

  1. A new context does not contain irrelevant information from the main session.
  2. The main context receives the result, not the complete work process.
  3. You can use a lower-cost model for a simple task.
  4. You can run independent tasks at the same time.

Always set a maximum number of subagents. Without this limit, Claude can start many subagents and quickly reach your plan limit.

Activity: Find the data structure

The file contains 300 unlabeled contacts. You must:

  • Find a common structure in the data.
  • Use that structure to extract data.
  • Check the quality of the extracted data.

You will make three outputs:

  • A common report schema.
  • An extraction skill.
  • A rubric for the extracted data.

A skill is a documented procedure that an LLM can follow. A rubric is a skill that contains criteria for judging LLM output.

This task requires repeated and careful work. The workflow that follows uses language models to complete the work quickly.

The pattern: subagents + adversarial review

First, lower-cost annotators extract labels from the data. An annotator is a subagent that extracts these labels. Then, an independent critic checks all labels and skills. The critic does not make labels. This separation of work is an adversarial review.

This pattern is similar to the independent review in plan, build, review. Use this pattern when:

  • You cannot read all the data.
  • You do not yet trust the extraction process.

Warning: This loop uses many tokens. Each cycle uses several annotators and one critic. Follow these group rules:

  • Select one student as the driver.
  • Only the driver runs the loop.
  • The driver shares their screen.
  • The other group members watch the loop.
  • Examine these intermediate outputs as a group:
    • The schema discussion.
    • The rubric.
    • The rejected labels.
  • Do not start a second set of subagents.

This is the only course activity that does not run separately in each repository.

Note: If you cannot continue, use the helper prompt to improve your prompt.

flowchart LR
    U[("contacts_unlabeled.jsonl")] --> H1 & H2 & H3
    H1["Haiku annotator 1"]
    H2["Haiku annotator 2"]
    H3["Haiku annotator n"]
    H1 & H2 & H3 -- "labels + skills" --> C["Sonnet critic: does not annotate"]
    C -- "rubric: accept/reject each label" --> G{"95% accepted, or 3 cycles?"}
    G -- "no: run again with the common skill" --> U
    G -- "yes" --> F[("accepted labels = intake fixtures")]
    C -. "proposed schema extensions" .-> S["your report schema"]
    H3 ~~~ S

Use these steps:

  1. Start the annotators. Ask Claude to start several Haiku subagents. Give each subagent a different part of contacts_unlabeled.jsonl. Also give each subagent your report schema. Each annotator must:

    • Independently extract all labels from its part of the data.
    • Write a skill that describes the procedure it used. The skill can contain mistakes.
  2. Run an independent review. Start one Sonnet critic that did not make labels. The critic reads every label and every skill. It returns:

    • A rubric with explicit accept or reject criteria for each label.
    • One common skill that combines the annotator skills. It keeps agreed procedures and resolves differences.
    • Proposed schema extensions for useful information that the current schema cannot store.
  3. Repeat the review cycle. Start new annotators that use the common skill. Then ask the critic to check the new labels.

The driver must tell an Opus agent to run this loop with /goal. The goal must include:

  • A success criterion of approximately 95% accepted labels.
  • A maximum of 3 cycles.
  • An instruction to write all intermediate results to disk.

The maximum prevents the goal from running without a limit. The group can examine the saved results. The driver can also give the results to the group.

This workflow has these important properties:

  • The lower-cost model reads every contact and creates the labels.
  • The higher-cost model reads the labels and skills. It uses its tokens for judgment instead of large-scale extraction.
  • You can examine the skill and rubric more quickly than hundreds of individual labels.
  • The independent review makes large-scale extraction more reliable.

The main lesson is simple: use a structured workflow to make model output more reliable.

When you are done

Only the driver ran the activity, but all group members need the outputs. You will use them in Stage 2: intake.

The driver must:

  1. Run npx ccusage.
  2. Take a screenshot of the result.
  3. Post the screenshot on Padlet. The result shows the cost of the group loop. Compare it with the other groups.
  4. Open a pull request in each teammate’s repository. Each pull request must contain:
    • The prompt that controlled the loop.
    • The common skill.
    • The rubric.
    • The accepted labels.

Each repository owner must review and merge their pull request. Do not approve it without reading it because you will use these outputs. If the critic proposed schema extensions, decide during the review whether to accept them.