ENGINEERING IN THE AGENT ERAFIELD NOTE 001

How I built an AI fleet and shipped Hurricane Harrison to the App Store.

How a growing library of proven prompts and repeatable workflows became Planestro, the system I built to coordinate an AI fleet while bringing Hurricane Harrison to the App Store.

Hurricane Harrison artwork: Harrison faces a green slime and a skeleton in a lantern-lit town beneath the game’s logo.

I started building Hurricane Harrison to explore new ways of making software with AI and to share a creative project with my son, the real Harrison. I began with one level and one gameplay loop.

As the game grew, so did the way I worked. I saved useful prompts, refined templates, and wrote down the steps that kept parallel agent sessions organized. That playbook became Planestro, the local orchestration system I built to coordinate a fleet of coding agents.

I still write the requirements, review plans, playtest every new feature across iPhone, iPad, and Mac, and decide what needs to change. Planestro automates the coordination I used to manage across all those sessions.

~170klines of code addedAcross Hurricane Harrison and Planestro, including tests and tools
43%higher delivery throughputMore Hurricane Harrison PRs merged than when I coordinated agents manually
57PRs improving Planestro itselfBuilding the system while using it to build the game
01 / THE ORIGIN

A game for Harrison.
A playbook for building it.

I brought the habits I’d developed over more than 20 years of building production software. When I approach a big problem, I break it into the smallest pieces that can still deliver working software. Each piece should build on a tested codebase and give someone something they can actually use.

For Hurricane Harrison, that meant starting with the epics needed to make a single playable level, then breaking each one into small deliverable tasks. I expected to work through them one at a time. Since this was a personal side project, a Markdown document was enough. I didn’t need to sign up for a separate tool to manage a Kanban board.

The document, HURRICANE_HARRISON_PROMPTS.md, became both my backlog and my working playbook. Each feature already had a description, so I added a subsection for the prompt I would use to ask an agent to plan it. I tracked its status and which model I intended to use.

The prompts became a process.

As I learned what worked, I saved successful prompts and developed templates for different kinds of work. A feature that needed new artwork required a Codex image-generation session first, followed by a Claude Code session to integrate the assets into the game.

I then found that the handoff worked better when Claude wrote the prompt for Codex and reviewed the plan Codex proposed before image generation began. I recorded that sequence in the document so I could repeat it on the next feature.

The templates kept evolving. I added instructions to give each session its own Git worktree, rebase with the latest develop before opening a pull request, and clean up the worktree and remote branch after the PR merged. Each addition captured something I wanted every session to do consistently.

By the time I retired the file, it ran to 2,458 lines. The same worktree instruction appeared 108 times.

FROM THE ORIGINAL PROMPT FILE
# MOBILE START SCREEN DOES NOT LOOK LIKE DESKTOP
## FABLE-5.1: FINISHED

Please perform this work in a new worktree
and add that to the plan.

start screen does not look quite right on my iPhone
versus how it looks on desktop

The playbook became software.

What I managed by hand What Planestro handles
Prompts with model headings and status notes Item files with planner, implementer, plan, and lifecycle state
Repeated instructions to create a worktree An isolated worktree for each worker item
My go-ahead to rebase, open a PR, and merge Deterministic integration and merge sequencing, with review holds
Reusable prompts for familiar kinds of work An editable library of reusable prompts and actions
Art first, then integration in another session Dependencies and epic branches coordinating mixed fleets
Reminders to remove branches and worktrees Cleanup gated by positive evidence that the work was merged

From one feature to parallel sessions.

As the workflow became more reliable, I could move faster. I put CI/CD in place to check new code against the requirements in each feature’s plan. For new gameplay mechanics, the plan also included a hold: the PR could not merge until I had played the build.

With those checks and review gates in place, I became comfortable working on several features at once. At times, I had five Claude Code sessions open alongside as many as five corresponding Codex image-generation sessions.

The playbook gave me confidence that familiar kinds of work would follow the same process and meet the same quality standards. But I was still coordinating the sessions myself. Each agent waited for my go-ahead before rebasing and creating its PR. Once the active PR merged, I let the next session update from develop and proceed. I also kept instructions for watching CI, fixing failures, and coming back to me when help was needed.

I was following the same patterns every time I started a feature. It became clear that much of this coordination could be automated, provided the system followed my playbook and kept me involved in the decisions and reviews that needed my judgment.

That became the next prompt in the same Markdown file: build a tool to run this workflow. The first prototype did not immediately replace my manual workflow. I kept coordinating game work by hand until Planestro was ready for sustained dogfooding.

“Use deterministic application logic wherever possible for Git, GitHub, worktree, process/session, CI status, persistence, cleanup, and state transitions.”The founding Planestro prompt

The prompt carried forward the approach I had developed by hand: use agents where reasoning is useful, put repeatable mechanics in application code, and preserve enough state to recover when a session disappears.

The original request already described more than launching agents: a visual backlog, a reusable prompt library, separate planner and implementer assignments in the proposed item format, and a lifecycle that distinguished merged code from finished cleanup. It required local, version-controlled planning files and recovery that checked what had actually happened in Git, GitHub, and running sessions.

It also described the workflow that became epics: generate art into a feature branch, let dependent code work follow it, then validate and merge the combined feature into develop. I asked for the tool’s source to live separately from the game so I could show the orchestration system without exposing the game’s source or planning data.

Read excerpts from the prompt that created Planestro

INTERNAL BACKLOG PLANNING & ORCHESTRATION TOOLING

Selected verbatim excerpts from the original request. Model names and wording are preserved. These are the founding requirements, not a description of every detail in the current implementation.

The boundary between code and judgment

Use deterministic application logic wherever possible for Git, GitHub, worktree, process/session, CI status, persistence, cleanup, and state transitions. The FABLE 5 orchestrating agent should be used where reasoning, planning, conflict assessment, prioritization, agent instruction, or other judgment is required rather than for mechanical polling or operations that can be reliably handled by application code.

Recovery was a requirement

All orchestration state must be durable and restart-safe. If the application or orchestrating agent stops unexpectedly, relaunching the tool must reconcile its persisted state against Git, local worktrees, running Claude sessions where recoverable, remote branches, GitHub pull requests, and CI rather than assuming previous operations completed.

The playbook remained readable

The tool also should have a configurable library of reusable prompts/actions for common operations such as creating a worktree, rebasing from develop, creating/merging a PR, monitoring CI, cleaning up a worktree/branch, cutting a release, auditing recent code, etc. These should remain available as human-readable version-controlled files and should be editable through the UI.

Art and code could depend on each other

Sometimes when writing out plans, I understand myself that a feature might be too big to do in one take and will have to break it up into multiple prompts. I might have a CODEX-5.6-SOL prompt that generates art and will have it create a pull request that targets a feature branch (e.g. feature/big_new_feature) and then I would like to have a claude agent that watches that remote branch for a new merged pull request so that it can start it's work on its own branch that targets that feature branch and once all of the chained dependent agents have merged their pull requests to that branch, the orchestrator can then rebase with the latest from develop and make sure CI is green before then merging the pull request and cleaning everything up.

More game work, while building the tool.

Using Planestro, I increased game delivery throughput by 43%: 124 merged Hurricane Harrison PRs compared with 87 while coordinating the agents myself, across equal ten-day windows. Alongside that game work, I merged 57 PRs improving Planestro itself. I was building the coordination system while using it to deliver more work.

Across the broader development period, about 170,000 lines were added across Hurricane Harrison and Planestro: roughly 139,000 lines of GDScript in the game repository and 29,000 lines of TypeScript in Planestro, including tests and tools.

Compared across two equal ten-day periods, counting merged Hurricane Harrison PRs and excluding epic integration merges and release promotions.

02 / THE OPERATING MODEL

Autonomy needs
an architecture.

I start with what I want and why. A planner reads the repository and proposes a plan. Plan approval is mine by default. I can explicitly delegate eligible plans to the orchestrator, with both a global setting and permission on the item. A worker then implements the approved plan in its own Git worktree.

The worker commits and reports through tools such as submit_plan, report_done, and report_blocked. Planestro owns the push, pull request, CI wait, rebase, and merge. Shared integration state stays behind an explicit boundary.

THE EPIC WORKFLOW Parallel work, shared destination

From an epic to an integrated feature.

PLANEpic plan and proposed breakdown

Define the feature, its child items, and their dependencies.

DAVIDI approve the breakdown

Creating the child work stays my decision.

Independent children can proceed in parallel
CHILD ACreate the artwork

Own worktree and PR into the epic branch.

CHILD C · DEPENDS ON AIntegrate the artwork

A separate item uses the merged assets.

CHILD B · INDEPENDENTBuild the game logic

Own worktree and PR; can proceed alongside the art work.

Child PR merges into the epic branch
FEATURE BRANCHThe children come together

Each child merges into the epic’s branch. A finished child advances the epic; it does not finish the whole feature.

VERIFY + REVIEWCheck the combined feature

Sync with develop, pass CI, and complete required review and playtesting.

PLANESTRO ENGINEMerge the epic into develop

Clean up only after confirming the work is safely merged.

Illustrative three-child epic. Actual breakdowns can contain many more children and dependencies.

Every worker child follows its own plan, approval, implementation, checks, and PR workflow. Dependencies control when work can begin; review holds control when it can merge.
Inside each worker child: plan, build, integrate, playtest

Plan

The planner investigates the repository and proposes the approach. Approval stays with me unless both global and per-item settings delegate an eligible plan. Epic breakdowns stay with me.

Build

The worker implements in its own worktree, runs the check suite, commits, and reports completion. A blocked worker reports the problem.

Integrate

Planestro pushes, opens the PR, waits for CI, syncs when the base moves, and merges after required gates. Epic children target the feature branch; standalone items target develop.

Playtest

I playtest every new feature across iPhone, iPad, and Mac and decide what needs to change. Required playtest and art reviews hold the relevant PR before it merges.

Large features become epics. Their children merge into a feature branch; the completed epic merges into develop. Art and code can move independently, then meet at a defined integration point.

The orchestrator watches the work; a facilitator steps in when asked.

THE PEOPLE, AGENTS, AND ENGINE Direction, execution, recovery
David

Writes the backlog, answers questions, plays builds.

↓ opens when needed
Facilitator agent · optional

Any Claude Code session I open to check on work or help a stuck orchestrator. Can start or stop Planestro, read logs and state, sync branches, and unblock items. Planestro runs independently; this session can come and go.

Planestro runs independently · its engine wakes the orchestrator
Orchestrator agent

Reviews eligible plans, rules on permissions, triages failures, and records feedback.

↔ David: questions and guidance
↓ instructions · ↑ reports, through tools
Child agents

Planners, workers, and Codex art sessions. Each item has its own worktree. These agents build the game.

Planestro engine · deterministic code

Under every layer: state, pushes, PRs, CI, merges, drains, and restarts. Agent requests to change orchestration state go through the engine.

I set direction and handle escalations. The facilitator is optional and on demand; the engine wakes the orchestrator to coordinate the children through Planestro’s tools.

I start Planestro from a terminal or any Claude Code session with ./start-planestro.sh, a wrapper over planestro start. The server runs detached in its own process session and keeps running after the launching session exits. ./stop-planestro.sh stops it. An optional facilitator can inspect it through the plugin’s /planestro command and read-only MCP tools.

The engine wakes the orchestrator on events. It rules on permission requests, investigates red CI, records decisions, and escalates questions. Its standing rules come from rulings I’ve already made.

That doesn’t remove the need for precise instructions. I once said “no generators,” meaning no power-generator objectives. It was relayed as a ban on layout generators across five epics. A supervisor can apply a decision consistently and still apply the wrong interpretation.

03 / INSIDE THE ENGINE

A state machine
under the AI fleet.

The agents can reason about the work. The engine decides which operations are legal, who can request them, and what evidence is required before the next step. That distinction is implemented in code.

Lifecycle and waiting are different things.

An item has one of nine lifecycle states, from NOT_YET_SUBMITTED to FINISHED. Each legal transition specifies both the actor and the item type: worker, epic, or external. A transition missing from the table is rejected.

SELECTED TRANSITIONS Authority is part of the rule
QUEUED
ENGINE ONLYWorker
PLAN_IN_PROGRESS
PLAN_READY
USER OR DELEGATED ORCHESTRATORWorker plan
IN_PROGRESS
PLAN_READY
USER ONLYEpic breakdown
IN_PROGRESS
IN_REVIEW
ENGINE ONLYMerge
MERGED

Waiting doesn’t erase that position. The item carries separate fields for blocked and its reason, awaiting, the current operation op, and triage. A plan waiting for approval remains PLAN_READY. Values such as “needs you” and the base branch are derived instead of stored as another potentially stale copy.

THE MODEL IN THE INTERFACE

The state stays put.
The reason stays visible.

HH-242 is ready for its epic breakdown to be approved. HH-240 is in review, waiting on CI. Both retain their lifecycle position while making the next dependency explicit.

Historical product capture. Its counts are separate from the October 1 metrics.

Approval is delegated explicitly.

Plan approval is mine by default. Delegation needs two switches: a global “May approve plans” setting and permission on the individual item. Even then, epic breakdowns, plans with a merge hold, art made outside Claude, and paid-tool budgets requiring human approval stay with me.

What the orchestrator may decide
Action Boundary
Approve a worker plan Both delegation settings must allow it; excluded plans return to the user.
Approve an epic breakdown User only. Creating child items commits more work and budget.
Request plan revisions Two fix rounds by default, then approve or hand over.
Release a review hold User only. The PR stays a draft and skips CI until approval.
Act after human takeover 27 action tools refuse. Reading and annotation remain available.
Merge Engine only, after integration checks pass.
The four plan decisions and the tool-level guards

A delegated review ends with approve_plan, amend_plan, request_plan_changes, or hand_plan_to_user. Amendments go back to the planner as stipulations so they become part of the revised plan. A wake without a decision hands the plan to me.

The engine enforces the configured fix-round limit. The worker-plan approval tool rejects plans flagged as epic-shaped. mark_blocked refuses while a turn or CI wait is running. These limits do not depend on the model remembering its instructions.

A ten-second tick moves the work.

The engine repeatedly runs a fixed pass: triage, disk checks, admission, reviews, epics, external work, integration, holds, stalls, and scheduled resumes. Draining prevents new work during restart preparation. Pausing stops the scheduling pass after the early checks.

Admission checks whether dependencies have merged, the epic is open, and a session slot is free. Slots count as occupied while something is actively driving the item.

See the engine tick from the source report
LiveEngine.ts · tick()
if (this.draining) return;
this.driveTriage();
void this.checkDisk();
if (this._paused) return;
this.admit();
this.nudgeStaleReviews();
this.driveStaleBlocks();
this.drivePlanReviews();
this.drivePermissionReviews();
await this.driveEpics();
await this.driveExternal();
this.driveInReview();
this.driveHeld();
await this.driveStalls();
this.driveResumeAfter();

CI waits beside the merge chain.

Syncs and merges run one at a time. A PR still waiting for CI leaves that serial chain with op: awaiting_ci, is polled outside it, and returns when green. A slow suite therefore doesn’t hold the integration lock while another PR is ready.

Before merging, the chain checks that the head is green on the current base tip. If the base moved, the item re-syncs. If it is already current, the engine can reuse its CI run. Git operations that update refs share a separate lock because worktrees share the same ref store; transient Git and GitHub errors are retried in the integration layer.

CI RETURNS A TYPED RESULT Nine possible verdicts
greenredblockedmovedtimeoutqueue_timeoutskippedstoppingtaken_over
How production incidents became CI rules
  • Runner queue timeout: queue_timeout distinguishes time waiting for a runner from a CI failure.
  • Stale PR head after a push: allow up to 90 seconds for propagation when GitHub still reports the old head. Other head changes yield moved.
  • A run fails without executing steps: inspect billing annotations and the short-run heuristic. A billing block is handled separately from a code failure, with a local-check fallback where configured.
  • The newest run remains cancelled: rerun once per head and give the replacement its own timeout budget.
  • Every job is skipped: resend ready_for_review once for the head, then stop if checks remain skipped. An intentional project path-filter exemption is handled separately and counts as green.
  • The same red run appears twice: do not count repeated observation as a second failed attempt.

The engine decides when judgment is needed.

The orchestrator doesn’t poll. Typed events wake it with context reconstructed from the item files: plan reviews, permissions, repeated CI failures, merges, blocks, and questions. The conversation can rotate every ten wakes because it isn’t the only place the decisions live.

Fourteen block reasons are routed for judgment, including sync conflicts, CI timeouts, and missing commits. Problems such as expired authentication, rate limits, or a full disk escalate directly to me.

FIRST RED

Back to the worker

The failure log goes to the session with the implementation context.

SECOND RED

Orchestrator triage

The item blocks as ci_red_twice and wakes the supervisor.

FOUR TRIAGE ROUNDS

Back to me

The cap applies per item per phase. The decision comes to a person.

The typed wake events

plan_review, permission_request, blocked_<reason>, ci_red_again, pr_opened, pr_merged, review_pending, orphans_found, scope_changed, decision_answered, user_message, epic_message, question_message, and question_review.

A human question doesn’t hold a process open.

When a command needs permission, Planestro records the request and stops that tool call. The worker is told to end its turn without retrying, working around the request, or reporting a block. The process exits with awaiting: permission persisted. The decision arrives as the first message of a resumed session.

There is no process waiting indefinitely for me to answer. The pending decision survives a restart.

DELEGATED APPROVAL, HUMAN JUDGMENT

“Have you played it?”

HH-288 can use delegated plan approval. It still asks whether I’ve played the finished Hometown and whether the mechanics are final enough to document.

The choices preserve partial answers: document one mechanic, wait on another, or hold the work until I’ve played. The approval checkbox doesn’t answer a product question on my behalf.

Cleanup needs fresh proof.

A remembered “merged” status isn’t enough to delete work. Cleanup fetches current refs, confirms the PR is merged, and checks that the worktree HEAD, local branch tip, and remote branch tip are ancestors of the target base. A dirty tree or unverifiable ancestry blocks cleanup.

See the cleanup proof from the source report
LiveEngine.ts · cleanup() · excerpt
await git.fetch(repo);
const view = await gh.prView(github, pr);
for (const [what, sha] of tips)
  if (!await git.isAncestor(repo, sha, `origin/${base}`))
    stranded.push(what);
if (view.state !== 'MERGED' || stranded.length)
  return block(id, 'cleanup_unverifiable',
    '… nothing was deleted');

The model gets bounded authority to make decisions. The engine owns the transition rules, gathers the evidence, and rejects operations outside those rules.

Why I built the coordination layer

Before committing to Planestro, I looked at existing tools for parallel coding agents, worktrees, task management, and integration. I found overlap with pieces of my workflow, but I didn’t find a tool that matched the whole process I wanted to run: backlog, planning, approval, model assignment, implementation, CI, merge sequencing, cleanup, and recovery.

That gave me a reason to build around established patterns. The engine reconciles recorded state with what has actually happened. A serial merge chain coordinates integration. Persisted plans and decisions let sessions stop and resume. The work was in connecting those mechanisms to the rules I had developed: which art a feature depends on, when a plan needs my judgment, and which changes must wait until I’ve played them.

I wanted the tool to work across projects from the start. Hurricane Harrison supplied the real work that shaped it. Planestro would own the coordination around the agents, while I kept responsibility for product direction and the decisions I chose to retain.

04 / THE QUALITY CONTRACT

Test what must stay true.

At the October 1 snapshot, Hurricane Harrison had 410 headless Godot checks. Workers can run them without opening a game window. CI uses them as an integration gate; held draft PRs skip runs until they’re ready.

For level geometry, a useful check pins the behavior I need. A lamp post must either intersect the wall or leave a full passage clear. Its coordinates can change. The player must still fit through.

ONE CONTRACT. THREE LAYOUTS.
✓ PASSIntersects the wall
✓ PASSFull passage clear
× FAILLooks open. Isn’t walkable.

Schematic geometry: gray = wall, blue = obstacle, dashed line = intended passage.

Mutation and negative tests deliberately introduce bad conditions to prove the checks catch them. When a playtest exposes a repeatable bug, I can turn that finding into a contract for future changes.

The suite still can’t decide whether a boss is fun. It can spawn, attack, and die correctly, then die in two seconds. That’s why I keep playing the builds.

What the quality records show

I brought the quality standard I’ve developed over more than twenty years of shipping production software. Plan approval, behavioral checks, review holds, and playtesting are how I put that standard into practice. The records give me a way to examine part of the result.

Of 194 Planestro-driven items whose PRs merged from September 16 through October 1, excluding epics, roughly four in five recorded no counted CI failures before merging.

Recorded CI failures per item Items Share
None 155 79.9%
One 29 14.9%
Two or more 10 5.2%

The first counted failure routes back to the worker with the implementation context; a second triggers orchestrator triage. These figures show how often merged work reached those failure thresholds. They do not measure defects after merge or capture every intervention along the way.

05 / THE FEEDBACK LOOP

The fleet started finding
bugs in its own tools.

One of the most useful things to watch has been the orchestrator reporting problems in Planestro itself. It has a dedicated note_planestro tool for missing capabilities, misfired rules, and friction in the workflow.

Those notes go to their own list on the Console, never into the game item. When the orchestrator is stuck, the facilitator steps in from outside Planestro. It is any Claude Code session I open when needed, with access to the shell, logs, and repository. Planestro keeps running independently when that session closes. It diagnoses the problem and files a GitHub issue named after the Hurricane Harrison item that exposed it.

Issue capture is automatic. The GitHub issues are filed whether or not a maintenance session is running. An optional workflow lets an active Claude Code session in the Planestro repository watch for issues and implement the fixes and enhancements I approve. If no session is running, the issues stay in GitHub and are suggested when the next session starts.

The repo agent writes the fix and tests, then takes it through a normal PR and green CI. I review and merge the PR, or approve the agent to merge it. The repo agent then pulls main into the running checkout. Code comments keep references to the incidents that prompted the changes. The 57 Planestro PRs merged alongside the game work show how much the tool was evolving; this feedback loop gave real development problems a route back into that work.

INCIDENT / HH-323 · HH-326 · HH-328OCT 1, 2026 · EDT

Three timeouts. One shared cause.

Cancelled CI runs had been requeued, but their timeout budgets still counted from the original wait. The orchestrator connected the failures and proposed resetting the clock.

The fix merged.

A replacement run now receives its full timeout. The code comment names all three HH items, preserving why the behavior exists.

The same day produced fixes for missing worker tools, a full disk, rerun controls, and CI queue handling. By the snapshot, the feedback list held 46 notes; 19 of Planestro’s 187 commits cited a Hurricane Harrison item.

I’m the escalation point, not a gate on every step. When the orchestrator is unsure, it asks me. I can fix the problem myself, but usually I guide it on how to clear the block. It adds the context the child needs, passes on the guidance, and follows the work through CI. The engine handles the merge once checks pass and any review I need to do, such as playtesting or reviewing art and screenshots, is complete. I approve which fixes and enhancements the optional maintenance workflow takes on. The repo agent carries that approved work through tests and green CI. I then review and merge, or approve the agent to merge.

The boundaries stay explicit. The orchestrator never edits Planestro. A separate repo agent makes the change through tests, a PR, and green CI. Before an update goes live, Planestro stops starting new work and waits for running turns, engine steps, and any in-flight orchestrator wake to finish. It then restarts on the new code, and the orchestrator resumes the children from persisted state. No child session holds the only copy of its work.

Planestro runs with --watch. After the repo agent pulls main, it drains running agent turns and restarts onto the updated code. The orchestrator then resumes the child items from saved state.

For work requiring a capability Planestro’s tools do not have, such as changing a CI workflow, using SSH, or handling a secret, the orchestrator can use file_for_maintainer. It creates a backlog item for the person running Planestro and never queues it for a worker.

Read the five incident-to-fix examples
  • 13:22 · Missing audio tool. HH-338’s plan needed ElevenLabs, but the worker lacked it. A capability check merged at 13:28 (#155), followed by a human-approved paid-tool budget rule at 13:49 (#157).
  • 15:04 · Full disk. HH-355 exposed an infrastructure problem as an item block. A guard to pause scheduling before the disk fills merged at 19:16 (#161).
  • 15:56 · Rerun policy. HH-326 remained constrained after disk space was freed. A user-tappable rerun card merged at 16:01 (#162).
  • 18:30 · Inherited timeout. HH-323, HH-326 and HH-328 exposed the old clock. The timeout reset merged at 18:36 (#164).
  • 21:18 · Queue backlog. The facilitator noted four queued items timing out. Queue-aware handling and admission control followed that evening (issue #170, PR #172).

Times are from the supplied journal and Git report. Issue records were often created alongside fixes; issue open-to-close time is not investigation duration.

06 / THE BOTTLENECK

More agents.
Same merge queue.

On October 1, I had 14 items in flight and two Mac runner instances sharing one suite lock. Concurrent suites dropped physics ticks, so verification was effectively serial.

The agents could keep producing changes. The verification system couldn’t keep up. Jobs waited for hours, and some timed out while still waiting for the lock.

14items in flight
1effective suite lane

October 1 configuration: two runner instances, one shared suite lock.

WHEN MORE AUTONOMY COSTS MORE

The recommendation was to stop spending.

For HH-247, the decision card recommends finishing by hand instead of resuming the worker. It also offers a budget increase or a hold, leaving the choice with me.

The displayed dollar figures are API-price estimates, not charges. The card records a recommendation and its rationale; it does not establish which option was chosen.

Earlier, a merge storm had exposed a different amplification problem: after every merge, Planestro re-pushed every open PR. A stale-head response could trigger another synchronization, and each merge spawned more CI work. A propagation grace period and avoiding unnecessary resyncs helped break that cycle.

Usage budget and playtest time were limits too. Rotating long orchestrator conversations and keeping state outside the conversation helped control context growth. Starting more agents was no longer the obvious way to ship faster.

The next step is scaling verification to match what the fleet can produce.

I built Planestro to standardize the way I prompted agents and remove opportunities for manual errors. Once it was working, it unlocked productivity I hadn’t realized was possible, while preserving the quality standards I expect from my products.

That increased throughput made verification capacity the next engineering problem to solve. Planestro now distinguishes time spent waiting in the CI queue from an actual test failure, and uses queue depth and admission control to manage incoming work. Those changes help the fleet operate within today’s limits. Expanding that capacity while preserving integration tests, review gates, and playtesting is the next stage of the work, and the subject of the third field note.

The first step: change-aware CI.

I have a plan for the next iteration; implementation hasn’t started yet. First, measure runner queue time, suite-lock waiting, and test execution separately, and fix the flaky checks that create avoidable reruns. Then map changes to the checks they can affect.

The selector will start in shadow mode: propose a smaller set of checks while the full suite still runs, so I can measure both missed failures and potential time savings. Targeted selection on epic branches comes only after that trial meets its safety and benefit gates. Unmapped or ambiguous changes fall back to the full suite, with full integration checks retained for gameplay changes reaching develop and main.

Test selection belongs in Hurricane Harrison’s CI workflow. Planestro stays project-neutral, recording validation evidence and making queue and execution times visible. The goal is to spend verification capacity where it matters while preserving the checks that protect the finished product.

07 / WHAT CARRIES FORWARD

The engineering is
in the boundaries.

Hurricane Harrison gave this work a demanding test bed: generated levels, physics timing, art handoffs, and one person accountable for the result. The questions carry into any team putting agents into a delivery pipeline.

Delivery is the outcome.

Hurricane Harrison shipped six App Store releases, from v1.0.0 through v1.5.0. The workflow carries changes from feature branches through develop and main, uses tags for versioning, and routes CI to hosted and self-hosted runners. Planestro grew around that delivery process.

One operator, by design

Planestro currently operates on behalf of one developer, coordinating agents across that developer’s projects. That is a deliberate design choice. Its plans, approvals, credentials, and operations all belong to one person who remains accountable for what the system does.

I’m preparing to open-source it because it has become useful enough in my own work that I want to share it with other engineers who might recognize the same need. They may already have a backlog and a good set of coding tools, but still spend considerable time coordinating the steps between deciding what to build and getting a tested change merged. That was the work my personal playbook captured, and it is the work Planestro now helps me manage.

Its backlog and planning features serve that purpose. They were never intended to replace Jira or Trello. If your team assigns you a Jira ticket, you can create a corresponding item in Planestro, or break it into several dependent items. Planestro helps you plan and carry out that work through implementation, review, CI, integration, and cleanup. Your team’s tracker remains where the shared requirements and priorities live.

That scope matters because Planestro can start agents and coordinate operations that change repositories, create pull requests, merge code, and remove worktrees and branches. Sharing one instance would require explicit rules about whose credentials each operation uses, who can approve which changes, and how one person’s agents are isolated from another person’s work. Switching from subscriptions to API keys would not, by itself, answer those questions.

The current design assumes one trusted operator. Other engineers can run their own instances for their own work, including work they contribute to a team. A shared instance would require a separate design for identity, permissions, and execution isolation. That remains outside the current scope while I focus on making the workflow I use reliable and useful to others.

Why this matters beyond games

For an engineering team, each mechanism answers a practical question about putting agents into production delivery.

Concern How Planestro handles it Why it matters to a team
Governance Plan approval starts with me. Delegation requires both the global setting and permission on the item; two fix rounds return the decision to a person. Review holds keep PRs in draft. Keep people involved where their judgment matters without making them approve every operation.
Least privilege Agents propose orchestration changes through narrow tools. The engine owns pushes, PRs, merges, and state; 27 action tools refuse requests on an item a person has taken over. Give agents useful work with enforceable limits on their authority.
Auditability Git tracks Markdown state with a semantic commit for each change, such as “HH-143: approve plan”. Reconstruct the decisions behind a change.
Recoverability Turns and wakes are rebuilt from persisted state, so sessions can be restarted or rotated. Keep work recoverable across crashes, dropped connections, and session replacement.
Safe operations Cleanup requires a fresh fetch, a merged PR, and ancestry checks on every tip. Dry-run mode and a global pause provide additional controls. Require positive evidence before destructive automation proceeds.
Quality Behavioral contracts are backed by mutation and negative tests. Repeatable playtest bugs become regression checks. Evaluate agent-written code against explicit behavior.
Cost control Per-turn cost tracking, usage and disk guards, orchestrator rotation, and branch-based runner routing make resource use visible and controllable. Understand what throughput costs and where resources are consumed.
Capacity Admission control responds to limits in CI, budget, and review capacity. Match the rate of new work to the system’s ability to verify and integrate it.

Planestro Design Principles

01

Put autonomy behind explicit boundaries.

Keep human gates where product judgment or approval matters. Let deterministic code own shared state, lifecycle transitions, and integration, with agents applying judgment through bounded tools.

02

Verify behavior, not implementation details.

Use contracts over coordinates and check the effective generated layout. Mutation tests deliberately introduce bad conditions to demonstrate that the checks catch them.

03

Turn product judgment into lasting requirements.

Playtesting produces rules the fleet can reuse. “Connected, never merely near” means a connection must exist in the generated geometry, not just look close. “Write Help last” keeps documentation tied to mechanics I have played and approved.

04

Park questions; release processes.

Persist the question, end the turn, and let the process exit. Resume with the answer when it arrives. Human review should not require an agent process to stay alive or a promise to remain pending.

05

Require proof before destructive cleanup.

A fresh fetch, a merged PR, and ancestry checks establish that the work is preserved before a worktree or branch is deleted. A dirty tree or ambiguous evidence blocks cleanup.

06

Keep recovery independent of agent memory.

No agent session holds the only copy of anything needed to recover. Every turn and orchestrator wake is rebuilt from persisted state, so sessions can be restarted or replaced. Files are memory. Conversations are cache.

07

Carry work through to delivery.

Coordinate implementation, review, CI, integration, release, and cleanup. Match new work to verification capacity, and turn incidents into tested improvements that help the next item through the pipeline.

This is the part of agentic engineering I find most interesting: designing the operating model that lets capable agents do useful work, while keeping responsibility for the outcome.

THE SERIES / BUILDING WITH AN AGENT FLEET

More from the workshop.

This article covers how Planestro came about, how the fleet operates, and how tests and supervision keep the work accountable. The follow-ups explore what happens as that workflow scales.

01 / YOU’RE HERE

Building an AI fleet

From a reusable prompt playbook to a working orchestration system: the operating model, deterministic engine, quality contracts, and agents supervising agents.

02 / PLANNED

Mixed fleets

How Codex art sessions and Claude Code workers contribute to the same feature, with explicit dependencies, review gates, and shared integration branches.

03 / PLANNED

When building gets faster than verification

Powerful agents and an efficient workflow can produce changes faster than a growing project can validate them. Integration tests, CI, reviews, and playtesting still take time and resources. How do you scale those checks, control their cost, and keep work moving while maintaining quality?

DG
BUILT AND WRITTEN BY

David Graves

I’m a software engineer with 20+ years building products across mobile, web, IoT, cloud, and AI. I care about the interfaces people use and the engineering systems that make them reliable.

I’m interested in Staff and Principal engineering roles where hands-on architecture, product judgment, and AI-enabled delivery meet.

See Hurricane Harrison on the App Store
Sources, scope, and what the numbers mean

This field note draws on project reports, original prompt excerpts, product captures, and journal/issue history supplied by David. September throughput counts were recalculated from the supplied per-PR data and checked against its raw Planestro PR snapshot. The engine claims were checked against Planestro source at commit ad3b68f (October 2, 2026), as documented in the supplied source review. The feedback-loop roles were supplied by David. Source-line figures come from the supplied recount.

  • Hurricane Harrison PRs exclude release promotions into main and feature-branch integration PRs that combine child PRs. Child PRs count once when merged, including merges into epic branches. Counts do not establish that every change reached develop or an App Store release during the same window.
  • September 6–15: 87 Hurricane Harrison PRs. September 16–25: 124, plus 57 Planestro PRs counted separately. The increase is 42.5%, rounded to 43%. September 1–15 provides a longer baseline of 107 PRs over 15 days; compared with 124 over ten days, the daily-rate increase is 73.8%.
  • The ~170k additions combine 138,653 GDScript lines added in Hurricane Harrison with 28,849 TypeScript lines added in Planestro, including tests and tools. The reporting windows differ: August 24–October 1 for the game and August 30–October 1 for Planestro. These are gross additions across both repositories, not net growth or code attributed solely to Planestro workers.
  • The ~27k source-line snapshot is 27,204 lines in Planestro as of October 1: 16,584 application source, 10,075 tests, and 545 site/prototype source. It includes comments and blank lines. It is a codebase-size snapshot, not September growth, code churn, or a productivity measure.
  • The 410 checks describe Hurricane Harrison’s suite at the snapshot. The feedback evidence records 46 notes from September 24–October 1. Notes, issues, commits, and fixed incidents are different measures.

The quality figures come from the supplied October 2 review of Hurricane Harrison item files and Git history: 194 merged, non-epic Planestro items during September 16–October 1, grouped by ci_failed_count. This is a merged-item failure-count sample, not a defect rate or a comparison with the manual workflow. The historical alternatives discussion reflects my account of looking for a workflow fit before committing to Planestro; it is not a current product comparison.

Engine source references · ad3b68f

  • packages/shared/src/transitions.ts
    TRANSITIONS: legal transitions, actors, and item kinds
  • packages/server/src/engine/LiveEngine.ts
    tick(); waitForCi() and its nine verdicts; HEAD_PROPAGATION_GRACE_MS; JUDGMENT_BLOCKS; MAX_TRIAGE_ROUNDS; cleanup()
  • packages/server/src/engine/orchestrator.ts
    HANDS_OFF: takeover restrictions on 27 tools; note_planestro
  • packages/shared/src/schema.ts
    may_approve_plans and max_plan_revisions: delegated approval and revision limits
  • packages/server/src/engine/sessions/LiveClaudePort.ts
    PARK_MSG: the permission-parking message
  • docs/operations.md
    “Restarting safely”: drain before restart