Leading Agent-Human Work Products

What a real curriculum production team taught me about running multi-agent orgs — a live task board, six work streams, and thirty published lesson documents, from a project that doubled as a pilot for discovering what a real agent-staffed training shop actually needs.

A student works at a laptop alongside classmates building and coding robotics projects at a STEM table.
Illustrative: the kind of hands-on STEM session this curriculum-production work is built for, week by week.
FIELD
REPORT
· NOT A
SIM ·
DRBILL360 — ORGANIZATIONAL DEVELOPMENT NOTES

Leading Agent-Human Work Products

What a real curriculum production team taught me about running multi-agent orgs — a live task board, six work streams, and thirty published lesson documents, from a project that doubled as a pilot for discovering what a real agent-staffed training shop actually needs.

Dr. William Hamilton, ODL Specialist — Training & eLearning August 2026 For anyone leading agent-human work products
From the board, as of this week
Leader W10–14 — 15 lessons
STREAM Adone
Builder W15–18 — 12 lessons, Doc IDs verified
STREAM Edone
Explorer W1 — 3 lessons
STREAM Ddone
Leader matrix regen — ~95 stale rows
STREAM Bneeds‑hamilton
Builder W19 — pending overlay decision
STREAM Enext
30
new lessons published
15
with a Doc ID rechecked before close-out
69
existing lessons audited, 0 needing rework
17
real mismatches surfaced for correction
Executive Summary

Dr. William Hamilton — an ODL specialist working in training and eLearning — spent roughly two weeks directing a team of AI agents to build curriculum for STEP: Skills, Training, Education, and Purpose. STEP is a youth leadership and mentoring initiative (STEP-Up) that moves youth through four progressive stages as they get older, called bands: Explorer, Builder, Leader, and Bridge.

The agents did not take instructions from Dr. Hamilton directly, task by task. They read a shared file called a task board — a single document that assigns work, tracks its status, and records what got done and how it was checked, often while Dr. Hamilton wasn’t watching in real time. Around that board, the work was split into streams: parallel roles, each with its own job description and its own explicit limits on what it was allowed to touch.

In that time, the team published 30 new lesson documents across three bands — Leader, Builder, and Explorer. Fifteen of those (Builder and Explorer) are logged on the board with a specific Drive document ID that was independently rechecked before being marked complete. Separately, a 69-lesson audit across two other bands confirmed that none of those existing lessons needed to be rewritten, while surfacing 17 real mismatches between the lessons and the program’s tracking matrix for later correction. One planned document edit was stopped by its own required test run before it could delete live content by mistake. One agent’s incorrect report of a technical failure was caught and corrected by that same agent before it ever reached Dr. Hamilton as a problem to solve.

None of this is a simulation, and none of it is a case study written after the fact — it is the actual, ongoing production record. It is also something else on purpose: because, to my knowledge, a team like this had never been assembled before, running it became a pilot for figuring out what roles, job descriptions, and rules a real agent-staffed training operation needs. That answer wasn’t planned in advance. It was discovered by doing the work.

The Organization

What the board is actually for

The board does two jobs, and both matter. Day to day, it’s how work gets assigned and completed: an item goes on the board, a stream picks it up when its turn comes, and when the work is done — a document, an audit finding, a fix — the result gets reported back onto the board, often after work that happened in the background, out of Dr. Hamilton’s direct view. That’s the operational job. The second job is just as real: because nothing like this team had existed before, running it revealed what it actually needs — which roles have to exist, what each role’s job description should say, what one role must never be allowed to do that another does routinely. The board isn’t only tracking curriculum work. It’s the record of an organization being designed in real time, by watching what breaks and what doesn’t.

next in-progress needs-hamilton done

Streams, and what they map to

“Stream” is the term used here on purpose — it captures both agents coordinating with each other and the human coordinating with all of them. Each stream also corresponds to a specific role with its own job description, the same relationship a functional team has to the position descriptions inside it:

StreamRole (job description)OwnsNever touches
ALesson Writer, Leader bandLeader Weeks 10–42, Worksheet PackMatrices, other bands
BMatrix AuditorObjective matrices, crosswalk auditsLesson prose (read-only role)
CDocs & PolicyManuals, governance, safety-sensitive editsLesson drafts, matrices
DLesson Writer, Explorer bandExplorer band (separate Drive tree)Other bands
ELesson Writer, Builder bandBuilder Weeks 15–42Weeks 1–14, closed under Stream B’s audit
MMechanical EditorTitles, links, formatting, any documentAny content judgment call

The rule that decides what’s safe to run at the same time

Two kinds of work happen on this board, and they get treated differently on purpose. Execution is work where the right answer has already been decided — applying an already-settled rule to one document after another. Discovery is work where the right answer isn’t known yet and requires a judgment call — deciding how an ambiguous term should be coded, or whether two documents that seem to conflict actually do. Execution is safe to hand to several streams at once. Discovery is not, and the reason is concrete: hand the same undecided judgment call to several agents running independently, and each one will very likely land on a different answer — each internally consistent, each looking correct on its own, because there’s no settled answer yet to check any of them against.

One real example from this project: a taxonomy correction that one agent worked through carefully produced seven fixes. Run the same undecided question across four independent agents, and the realistic outcome isn’t seven correct fixes reached four times — it’s four different sets of guesses, each self-consistent, most of them wrong, discovered only much later once other lessons had already been built on top of the wrong ones. So the rule is: decide the judgment call once, with one agent or the human doing the deciding, and only fan the decided answer out to multiple streams once there’s a single answer to check against.

The hard stop — and why it’s really an HR and Operations decision-rights policy

A small set of item types on the board carries a status called needs-hamilton, and nothing downstream proceeds while an item sits in that state — no stream acts on a guess about how Dr. Hamilton will rule. What lands on that list isn’t arbitrary: real content changes to safety-sensitive material, taxonomy changes that ripple across the whole coding system, anything a family or funder will see, anything that would overwrite or delete original content instead of creating a new version. This is the same principle any real HR or Ops shop already runs on: most day-to-day decisions get delegated to staff, but a specific, named list of decisions is reserved for a manager or director regardless of who’s doing the work — not because the staff aren’t capable, but because those decisions are expensive or impossible to reverse, and someone has to be accountable for them by design.

Two real instances from this project: the Leader objective matrix regeneration sat at needs-hamilton, evidence pack fully built, waiting on Dr. Hamilton’s scope decision before any stream touched it. A Builder-band lesson update sat unfinished for the same reason — the stream had the content it needed but stopped and asked before applying it to a live lesson.

Practical note This is what “building the organization” actually looked like here — not writing an org chart first, but running real work and writing down, after each incident, exactly what decision needed to be reserved and why.
What Held Under Real Production Pressure

A caught bug — walked through in full, because the details are the point

The plan: delete three sections from a live program manual that had accidentally been duplicated — a manual staff actually use, not a draft.

Why a dry run came first: a dry run is a real call to the editing tool that reports exactly what it would delete — the actual text, the actual paragraph count — without deleting anything. It’s required before any live document deletion because these documents can’t be casually undone; nothing gets deleted for real until a rehearsal shows the rehearsal did what was intended.

What “anchor” means here: the deletion tool doesn’t work by line numbers. It works by anchors — a piece of text you point to as the start of what to delete, and a piece of text you point to as the stop. Think of it as “cut from this landmark to that landmark” rather than “cut lines 40 through 90.” For each of the three deletions, the natural stop-anchor was the heading of the next section — the obvious landmark to use.

What the dry run caught: the tool’s delete range turned out to be “end-inclusive” — it deletes everything up to and including the stop-anchor itself, not up to it and stopping short. Used as planned, all three deletions would have deleted the next section’s live heading along with the intended duplicate. Real, currently-used content would have been lost, not just the target.

What happened next: the stream stopped. It deleted nothing, guessed at nothing, and reported exactly what it found and why it couldn’t proceed as planned. The Coordinator rewrote the three stop-anchors to point at the actual last line of each duplicate section instead of the next section’s heading — then “re-verified” them, meaning: checked each new start-and-stop pair against the live document a second time, confirming each one pointed at exactly the intended text and nothing else, for all three deletions separately, before running anything for real.

How the result was confirmed, not just trusted: after the real deletion ran, the document’s total paragraph count was checked against a number worked out in advance — the original count, minus exactly what the three deletions were supposed to remove. The two numbers matched exactly. That match is what confirms nothing extra was lost and nothing intended was left behind — an independent count, not a trust in the tool’s own “it worked” response.

So what: the required rehearsal step caught a real bug in the editing tool before it could damage a document real staff use — not because an agent happened to be careful that day, but because “dry run before any live document edit” is a rule that doesn’t depend on anyone remembering to be careful.

A self-corrected false alarm

What happened: a stream reported that it couldn’t open a source document it needed. That’s normally read as a permissions problem — “this document isn’t shared with the right account” — the kind of thing that would ordinarily need Dr. Hamilton to go check a sharing setting. Before that report went anywhere, the same stream re-checked its own request in the same session and found the real cause: an earlier document listing had truncated the file’s ID to 24 characters for display, and the stream had used that shortened ID as if it were the real one. It wasn’t being denied access — it was looking up a document that, as ID’d, didn’t fully exist.

Why it matters: had the first report gone to Dr. Hamilton as originally stated, the likely response would have been to go check or change sharing settings on a document that was never actually restricted — real time spent solving a problem that didn’t exist, while the real fix, using the complete ID, sat unfound.

So what: the stream didn’t treat its own first read of the failure as fact. It re-checked before escalating, found the actual cause, fixed the immediate case, and wrote the root cause onto the board so the next session wouldn’t repeat it. That’s the “verify before you trust a report” habit running in both directions — checking another agent’s report, and an agent checking its own.

This file does not make an agent self-starting. It tells a worker what to do once it’s running — someone still has to open the terminal and start it.
Key Takeaways
  • The board runs two jobs at once: it executes today’s work, and it’s the running record of an organization figuring out what roles and rules it needs.
  • “Stream” maps directly to a role and a job description — an actual position with an owns/never-touches boundary, not an abstraction.
  • Separate “discovery” (an undecided judgment call) from “execution” (applying a decision already made) before deciding what’s safe to run in parallel — parallel discovery produces several confident wrong answers, not a faster right one.
  • Keep a short, named list of decisions for the human only — the ones too costly or risky to reverse. That’s the same thing any HR or Ops shop does by reserving certain calls for a manager, and it’s what makes delegating everything else safe.
  • Required rehearsal steps (dry runs) and independent after-the-fact checks (recounts, re-verification) catch real bugs before they become real damage, without depending on anyone remembering to be careful.
  • Verification runs in both directions: checking another agent’s report, and an agent checking its own first read of a problem before escalating it.
Next Step

Don’t try to design the full organization before starting. Start with a starter board — the roles and rules you already know you’ll need, nothing more — and expect the real work to surface the rest: which additional streams, roles, and job descriptions you actually need, discovered the same way this project found its own.

Drawn from a live production log and task board kept alongside STEP — Skills, Training, Education, and Purpose — the youth leadership and mentoring initiative (STEP-Up) currently in production. Not a simulation, not a case study written after the fact.

Leave a Reply

Your email address will not be published. Required fields are marked *