A game for Harrison.
A playbook for building it.
I brought the habits I’d developed over more than 20 years of building production software. When I approach a big problem, I break it into the smallest pieces that can still deliver working software. Each piece should build on a tested codebase and give someone something they can actually use.
For Hurricane Harrison, that meant starting with the epics needed to make a single playable level, then breaking each one into small deliverable tasks. I expected to work through them one at a time. Since this was a personal side project, a Markdown document was enough. I didn’t need to sign up for a separate tool to manage a Kanban board.
The document, HURRICANE_HARRISON_PROMPTS.md, became
both my backlog and my working playbook. Each feature already had
a description, so I added a subsection for the prompt I would use
to ask an agent to plan it. I tracked its status and which model I
intended to use.
The prompts became a process.
As I learned what worked, I saved successful prompts and developed templates for different kinds of work. A feature that needed new artwork required a Codex image-generation session first, followed by a Claude Code session to integrate the assets into the game.
I then found that the handoff worked better when Claude wrote the prompt for Codex and reviewed the plan Codex proposed before image generation began. I recorded that sequence in the document so I could repeat it on the next feature.
The templates kept evolving. I added instructions to give each
session its own Git worktree, rebase with the latest
develop before opening a pull request, and clean up
the worktree and remote branch after the PR merged. Each addition
captured something I wanted every session to do consistently.
By the time I retired the file, it ran to 2,458 lines. The same worktree instruction appeared 108 times.
# MOBILE START SCREEN DOES NOT LOOK LIKE DESKTOP ## FABLE-5.1: FINISHED Please perform this work in a new worktree and add that to the plan. start screen does not look quite right on my iPhone versus how it looks on desktop
The playbook became software.
| What I managed by hand | What Planestro handles |
|---|---|
| Prompts with model headings and status notes | Item files with planner, implementer, plan, and lifecycle state |
| Repeated instructions to create a worktree | An isolated worktree for each worker item |
| My go-ahead to rebase, open a PR, and merge | Deterministic integration and merge sequencing, with review holds |
| Reusable prompts for familiar kinds of work | An editable library of reusable prompts and actions |
| Art first, then integration in another session | Dependencies and epic branches coordinating mixed fleets |
| Reminders to remove branches and worktrees | Cleanup gated by positive evidence that the work was merged |
From one feature to parallel sessions.
As the workflow became more reliable, I could move faster. I put CI/CD in place to check new code against the requirements in each feature’s plan. For new gameplay mechanics, the plan also included a hold: the PR could not merge until I had played the build.
With those checks and review gates in place, I became comfortable working on several features at once. At times, I had five Claude Code sessions open alongside as many as five corresponding Codex image-generation sessions.
The playbook gave me confidence that familiar kinds of work would
follow the same process and meet the same quality standards. But I
was still coordinating the sessions myself. Each agent waited for
my go-ahead before rebasing and creating its PR. Once the active
PR merged, I let the next session update from
develop and proceed. I also kept instructions for
watching CI, fixing failures, and coming back to me when help was
needed.
I was following the same patterns every time I started a feature. It became clear that much of this coordination could be automated, provided the system followed my playbook and kept me involved in the decisions and reviews that needed my judgment.
That became the next prompt in the same Markdown file: build a tool to run this workflow. The first prototype did not immediately replace my manual workflow. I kept coordinating game work by hand until Planestro was ready for sustained dogfooding.
“Use deterministic application logic wherever possible for Git, GitHub, worktree, process/session, CI status, persistence, cleanup, and state transitions.”The founding Planestro prompt
The prompt carried forward the approach I had developed by hand: use agents where reasoning is useful, put repeatable mechanics in application code, and preserve enough state to recover when a session disappears.
The original request already described more than launching agents: a visual backlog, a reusable prompt library, separate planner and implementer assignments in the proposed item format, and a lifecycle that distinguished merged code from finished cleanup. It required local, version-controlled planning files and recovery that checked what had actually happened in Git, GitHub, and running sessions.
It also described the workflow that became epics: generate art
into a feature branch, let dependent code work follow it, then
validate and merge the combined feature into develop.
I asked for the tool’s source to live separately from the game so
I could show the orchestration system without exposing the game’s
source or planning data.
Read excerpts from the prompt that created Planestro
INTERNAL BACKLOG PLANNING & ORCHESTRATION TOOLING
Selected verbatim excerpts from the original request. Model names and wording are preserved. These are the founding requirements, not a description of every detail in the current implementation.
The boundary between code and judgment
Use deterministic application logic wherever possible for Git, GitHub, worktree, process/session, CI status, persistence, cleanup, and state transitions. The FABLE 5 orchestrating agent should be used where reasoning, planning, conflict assessment, prioritization, agent instruction, or other judgment is required rather than for mechanical polling or operations that can be reliably handled by application code.
Recovery was a requirement
All orchestration state must be durable and restart-safe. If the application or orchestrating agent stops unexpectedly, relaunching the tool must reconcile its persisted state against Git, local worktrees, running Claude sessions where recoverable, remote branches, GitHub pull requests, and CI rather than assuming previous operations completed.
The playbook remained readable
The tool also should have a configurable library of reusable prompts/actions for common operations such as creating a worktree, rebasing from develop, creating/merging a PR, monitoring CI, cleaning up a worktree/branch, cutting a release, auditing recent code, etc. These should remain available as human-readable version-controlled files and should be editable through the UI.
Art and code could depend on each other
Sometimes when writing out plans, I understand myself that a feature might be too big to do in one take and will have to break it up into multiple prompts. I might have a CODEX-5.6-SOL prompt that generates art and will have it create a pull request that targets a feature branch (e.g. feature/big_new_feature) and then I would like to have a claude agent that watches that remote branch for a new merged pull request so that it can start it's work on its own branch that targets that feature branch and once all of the chained dependent agents have merged their pull requests to that branch, the orchestrator can then rebase with the latest from develop and make sure CI is green before then merging the pull request and cleaning everything up.
More game work, while building the tool.
Using Planestro, I increased game delivery throughput by 43%: 124 merged Hurricane Harrison PRs compared with 87 while coordinating the agents myself, across equal ten-day windows. Alongside that game work, I merged 57 PRs improving Planestro itself. I was building the coordination system while using it to deliver more work.
Across the broader development period, about 170,000 lines were added across Hurricane Harrison and Planestro: roughly 139,000 lines of GDScript in the game repository and 29,000 lines of TypeScript in Planestro, including tests and tools.
Compared across two equal ten-day periods, counting merged Hurricane Harrison PRs and excluding epic integration merges and release promotions.
Autonomy needs
an architecture.
I start with what I want and why. A planner reads the repository and proposes a plan. Plan approval is mine by default. I can explicitly delegate eligible plans to the orchestrator, with both a global setting and permission on the item. A worker then implements the approved plan in its own Git worktree.
The worker commits and reports through tools such as
submit_plan, report_done, and
report_blocked. Planestro owns the push, pull
request, CI wait, rebase, and merge. Shared integration state
stays behind an explicit boundary.
From an epic to an integrated feature.
Define the feature, its child items, and their dependencies.
Creating the child work stays my decision.
Own worktree and PR into the epic branch.
A separate item uses the merged assets.
Own worktree and PR; can proceed alongside the art work.
Each child merges into the epic’s branch. A finished child advances the epic; it does not finish the whole feature.
Sync with develop, pass CI, and complete required review and playtesting.
Clean up only after confirming the work is safely merged.
Illustrative three-child epic. Actual breakdowns can contain many more children and dependencies.
Inside each worker child: plan, build, integrate, playtest
Plan
The planner investigates the repository and proposes the approach. Approval stays with me unless both global and per-item settings delegate an eligible plan. Epic breakdowns stay with me.
Build
The worker implements in its own worktree, runs the check suite, commits, and reports completion. A blocked worker reports the problem.
Integrate
Planestro pushes, opens the PR, waits for CI, syncs when the base moves, and merges after required gates. Epic children target the feature branch; standalone items target develop.
Playtest
I playtest every new feature across iPhone, iPad, and Mac and decide what needs to change. Required playtest and art reviews hold the relevant PR before it merges.
Large features become epics. Their children merge into a feature
branch; the completed epic merges into develop. Art
and code can move independently, then meet at a defined
integration point.
The orchestrator watches the work; a facilitator steps in when asked.
Writes the backlog, answers questions, plays builds.
Any Claude Code session I open to check on work or help a stuck orchestrator. Can start or stop Planestro, read logs and state, sync branches, and unblock items. Planestro runs independently; this session can come and go.
Reviews eligible plans, rules on permissions, triages failures, and records feedback.
↔ David: questions and guidancePlanners, workers, and Codex art sessions. Each item has its own worktree. These agents build the game.
Under every layer: state, pushes, PRs, CI, merges, drains, and restarts. Agent requests to change orchestration state go through the engine.
I start Planestro from a terminal or any Claude Code session with
./start-planestro.sh, a wrapper over
planestro start. The server runs detached in its own
process session and keeps running after the launching session
exits. ./stop-planestro.sh stops it. An optional
facilitator can inspect it through the plugin’s
/planestro command and read-only MCP tools.
The engine wakes the orchestrator on events. It rules on permission requests, investigates red CI, records decisions, and escalates questions. Its standing rules come from rulings I’ve already made.
That doesn’t remove the need for precise instructions. I once said “no generators,” meaning no power-generator objectives. It was relayed as a ban on layout generators across five epics. A supervisor can apply a decision consistently and still apply the wrong interpretation.
The actual working surface.
Product captures · October 1Tap a capture to inspect it. The board’s 302 completed items are a later snapshot than the 300 in the source metrics.
A state machine
under the AI fleet.
The agents can reason about the work. The engine decides which operations are legal, who can request them, and what evidence is required before the next step. That distinction is implemented in code.
Lifecycle and waiting are different things.
An item has one of nine lifecycle states, from
NOT_YET_SUBMITTED to FINISHED. Each
legal transition specifies both the actor and the item type:
worker, epic, or external. A transition missing from the table is
rejected.
QUEUED
PLAN_IN_PROGRESS
PLAN_READY
IN_PROGRESS
PLAN_READY
IN_PROGRESS
IN_REVIEW
MERGED
Waiting doesn’t erase that position. The item carries separate
fields for blocked and its reason,
awaiting, the current operation op, and
triage. A plan waiting for approval remains
PLAN_READY. Values such as “needs you” and the base
branch are derived instead of stored as another potentially stale
copy.
The state stays put.
The reason stays visible.
HH-242 is ready for its epic breakdown to be approved. HH-240 is in review, waiting on CI. Both retain their lifecycle position while making the next dependency explicit.
Historical product capture. Its counts are separate from the October 1 metrics.
Approval is delegated explicitly.
Plan approval is mine by default. Delegation needs two switches: a global “May approve plans” setting and permission on the individual item. Even then, epic breakdowns, plans with a merge hold, art made outside Claude, and paid-tool budgets requiring human approval stay with me.
The four plan decisions and the tool-level guards
A delegated review ends with approve_plan,
amend_plan, request_plan_changes, or
hand_plan_to_user. Amendments go back to the
planner as stipulations so they become part of the revised
plan. A wake without a decision hands the plan to me.
The engine enforces the configured fix-round limit. The
worker-plan approval tool rejects plans flagged as
epic-shaped. mark_blocked refuses while a turn or
CI wait is running. These limits do not depend on the model
remembering its instructions.
A ten-second tick moves the work.
The engine repeatedly runs a fixed pass: triage, disk checks, admission, reviews, epics, external work, integration, holds, stalls, and scheduled resumes. Draining prevents new work during restart preparation. Pausing stops the scheduling pass after the early checks.
Admission checks whether dependencies have merged, the epic is open, and a session slot is free. Slots count as occupied while something is actively driving the item.
See the engine tick from the source report
if (this.draining) return; this.driveTriage(); void this.checkDisk(); if (this._paused) return; this.admit(); this.nudgeStaleReviews(); this.driveStaleBlocks(); this.drivePlanReviews(); this.drivePermissionReviews(); await this.driveEpics(); await this.driveExternal(); this.driveInReview(); this.driveHeld(); await this.driveStalls(); this.driveResumeAfter();
CI waits beside the merge chain.
Syncs and merges run one at a time. A PR still waiting for CI
leaves that serial chain with op: awaiting_ci, is
polled outside it, and returns when green. A slow suite therefore
doesn’t hold the integration lock while another PR is ready.
Before merging, the chain checks that the head is green on the current base tip. If the base moved, the item re-syncs. If it is already current, the engine can reuse its CI run. Git operations that update refs share a separate lock because worktrees share the same ref store; transient Git and GitHub errors are retried in the integration layer.
greenredblockedmovedtimeoutqueue_timeoutskippedstoppingtaken_over
How production incidents became CI rules
-
Runner queue timeout:
queue_timeoutdistinguishes time waiting for a runner from a CI failure. -
Stale PR head after a push: allow up to 90
seconds for propagation when GitHub still reports the old
head. Other head changes yield
moved. - A run fails without executing steps: inspect billing annotations and the short-run heuristic. A billing block is handled separately from a code failure, with a local-check fallback where configured.
- The newest run remains cancelled: rerun once per head and give the replacement its own timeout budget.
-
Every job is skipped: resend
ready_for_reviewonce for the head, then stop if checks remain skipped. An intentional project path-filter exemption is handled separately and counts as green. - The same red run appears twice: do not count repeated observation as a second failed attempt.
The engine decides when judgment is needed.
The orchestrator doesn’t poll. Typed events wake it with context reconstructed from the item files: plan reviews, permissions, repeated CI failures, merges, blocks, and questions. The conversation can rotate every ten wakes because it isn’t the only place the decisions live.
Fourteen block reasons are routed for judgment, including sync conflicts, CI timeouts, and missing commits. Problems such as expired authentication, rate limits, or a full disk escalate directly to me.
Back to the worker
The failure log goes to the session with the implementation context.
Orchestrator triage
The item blocks as ci_red_twice and wakes the
supervisor.
Back to me
The cap applies per item per phase. The decision comes to a person.
The typed wake events
plan_review, permission_request,
blocked_<reason>,
ci_red_again, pr_opened,
pr_merged, review_pending,
orphans_found, scope_changed,
decision_answered, user_message,
epic_message, question_message, and
question_review.
A human question doesn’t hold a process open.
When a command needs permission, Planestro records the request and
stops that tool call. The worker is told to end its turn without
retrying, working around the request, or reporting a block. The
process exits with awaiting: permission persisted.
The decision arrives as the first message of a resumed session.
There is no process waiting indefinitely for me to answer. The pending decision survives a restart.
“Have you played it?”
HH-288 can use delegated plan approval. It still asks whether I’ve played the finished Hometown and whether the mechanics are final enough to document.
The choices preserve partial answers: document one mechanic, wait on another, or hold the work until I’ve played. The approval checkbox doesn’t answer a product question on my behalf.
Cleanup needs fresh proof.
A remembered “merged” status isn’t enough to delete work. Cleanup fetches current refs, confirms the PR is merged, and checks that the worktree HEAD, local branch tip, and remote branch tip are ancestors of the target base. A dirty tree or unverifiable ancestry blocks cleanup.
See the cleanup proof from the source report
await git.fetch(repo);
const view = await gh.prView(github, pr);
for (const [what, sha] of tips)
if (!await git.isAncestor(repo, sha, `origin/${base}`))
stranded.push(what);
if (view.state !== 'MERGED' || stranded.length)
return block(id, 'cleanup_unverifiable',
'… nothing was deleted');
The model gets bounded authority to make decisions. The engine owns the transition rules, gathers the evidence, and rejects operations outside those rules.
Why I built the coordination layer
Before committing to Planestro, I looked at existing tools for parallel coding agents, worktrees, task management, and integration. I found overlap with pieces of my workflow, but I didn’t find a tool that matched the whole process I wanted to run: backlog, planning, approval, model assignment, implementation, CI, merge sequencing, cleanup, and recovery.
That gave me a reason to build around established patterns. The engine reconciles recorded state with what has actually happened. A serial merge chain coordinates integration. Persisted plans and decisions let sessions stop and resume. The work was in connecting those mechanisms to the rules I had developed: which art a feature depends on, when a plan needs my judgment, and which changes must wait until I’ve played them.
I wanted the tool to work across projects from the start. Hurricane Harrison supplied the real work that shaped it. Planestro would own the coordination around the agents, while I kept responsibility for product direction and the decisions I chose to retain.
Test what must stay true.
At the October 1 snapshot, Hurricane Harrison had 410 headless Godot checks. Workers can run them without opening a game window. CI uses them as an integration gate; held draft PRs skip runs until they’re ready.
For level geometry, a useful check pins the behavior I need. A lamp post must either intersect the wall or leave a full passage clear. Its coordinates can change. The player must still fit through.
Schematic geometry: gray = wall, blue = obstacle, dashed line = intended passage.
Mutation and negative tests deliberately introduce bad conditions to prove the checks catch them. When a playtest exposes a repeatable bug, I can turn that finding into a contract for future changes.
The suite still can’t decide whether a boss is fun. It can spawn, attack, and die correctly, then die in two seconds. That’s why I keep playing the builds.
What the quality records show
I brought the quality standard I’ve developed over more than twenty years of shipping production software. Plan approval, behavioral checks, review holds, and playtesting are how I put that standard into practice. The records give me a way to examine part of the result.
Of 194 Planestro-driven items whose PRs merged from September 16 through October 1, excluding epics, roughly four in five recorded no counted CI failures before merging.
| Recorded CI failures per item | Items | Share |
|---|---|---|
| None | 155 | 79.9% |
| One | 29 | 14.9% |
| Two or more | 10 | 5.2% |
The first counted failure routes back to the worker with the implementation context; a second triggers orchestrator triage. These figures show how often merged work reached those failure thresholds. They do not measure defects after merge or capture every intervention along the way.
The fleet started finding
bugs in its own tools.
One of the most useful things to watch has been the orchestrator
reporting problems in Planestro itself. It has a dedicated
note_planestro tool for missing capabilities,
misfired rules, and friction in the workflow.
Those notes go to their own list on the Console, never into the game item. When the orchestrator is stuck, the facilitator steps in from outside Planestro. It is any Claude Code session I open when needed, with access to the shell, logs, and repository. Planestro keeps running independently when that session closes. It diagnoses the problem and files a GitHub issue named after the Hurricane Harrison item that exposed it.
Issue capture is automatic. The GitHub issues are filed whether or not a maintenance session is running. An optional workflow lets an active Claude Code session in the Planestro repository watch for issues and implement the fixes and enhancements I approve. If no session is running, the issues stay in GitHub and are suggested when the next session starts.
The repo agent writes the fix and tests, then takes it through a
normal PR and green CI. I review and merge the PR, or approve the
agent to merge it. The repo agent then pulls
main into the running checkout. Code comments keep
references to the incidents that prompted the changes. The 57
Planestro PRs merged alongside the game work show how much the
tool was evolving; this feedback loop gave real development
problems a route back into that work.
Three timeouts. One shared cause.
Cancelled CI runs had been requeued, but their timeout budgets still counted from the original wait. The orchestrator connected the failures and proposed resetting the clock.
The fix merged.
A replacement run now receives its full timeout. The code comment names all three HH items, preserving why the behavior exists.
The orchestrator notes the flaw; the facilitator investigates and files an issue.
An active repo session opens a PR with tests. After green CI, I merge or approve the agent to merge. Without an active session, the issue waits.
The repo agent pulls main. Planestro’s watch mode drains and restarts; the orchestrator resumes the children from saved state.
The same day produced fixes for missing worker tools, a full disk, rerun controls, and CI queue handling. By the snapshot, the feedback list held 46 notes; 19 of Planestro’s 187 commits cited a Hurricane Harrison item.
I’m the escalation point, not a gate on every step. When the orchestrator is unsure, it asks me. I can fix the problem myself, but usually I guide it on how to clear the block. It adds the context the child needs, passes on the guidance, and follows the work through CI. The engine handles the merge once checks pass and any review I need to do, such as playtesting or reviewing art and screenshots, is complete. I approve which fixes and enhancements the optional maintenance workflow takes on. The repo agent carries that approved work through tests and green CI. I then review and merge, or approve the agent to merge.
The boundaries stay explicit. The orchestrator never edits Planestro. A separate repo agent makes the change through tests, a PR, and green CI. Before an update goes live, Planestro stops starting new work and waits for running turns, engine steps, and any in-flight orchestrator wake to finish. It then restarts on the new code, and the orchestrator resumes the children from persisted state. No child session holds the only copy of its work.
Planestro runs with --watch. After the repo agent
pulls main, it drains running agent turns and
restarts onto the updated code. The orchestrator then resumes the
child items from saved state.
For work requiring a capability Planestro’s tools do not have,
such as changing a CI workflow, using SSH, or handling a secret,
the orchestrator can use file_for_maintainer. It
creates a backlog item for the person running Planestro and never
queues it for a worker.
Read the five incident-to-fix examples
- 13:22 · Missing audio tool. HH-338’s plan needed ElevenLabs, but the worker lacked it. A capability check merged at 13:28 (#155), followed by a human-approved paid-tool budget rule at 13:49 (#157).
- 15:04 · Full disk. HH-355 exposed an infrastructure problem as an item block. A guard to pause scheduling before the disk fills merged at 19:16 (#161).
- 15:56 · Rerun policy. HH-326 remained constrained after disk space was freed. A user-tappable rerun card merged at 16:01 (#162).
- 18:30 · Inherited timeout. HH-323, HH-326 and HH-328 exposed the old clock. The timeout reset merged at 18:36 (#164).
- 21:18 · Queue backlog. The facilitator noted four queued items timing out. Queue-aware handling and admission control followed that evening (issue #170, PR #172).
Times are from the supplied journal and Git report. Issue records were often created alongside fixes; issue open-to-close time is not investigation duration.
More agents.
Same merge queue.
On October 1, I had 14 items in flight and two Mac runner instances sharing one suite lock. Concurrent suites dropped physics ticks, so verification was effectively serial.
The agents could keep producing changes. The verification system couldn’t keep up. Jobs waited for hours, and some timed out while still waiting for the lock.
October 1 configuration: two runner instances, one shared suite lock.
The recommendation was to stop spending.
For HH-247, the decision card recommends finishing by hand instead of resuming the worker. It also offers a budget increase or a hold, leaving the choice with me.
The displayed dollar figures are API-price estimates, not charges. The card records a recommendation and its rationale; it does not establish which option was chosen.
Earlier, a merge storm had exposed a different amplification problem: after every merge, Planestro re-pushed every open PR. A stale-head response could trigger another synchronization, and each merge spawned more CI work. A propagation grace period and avoiding unnecessary resyncs helped break that cycle.
Usage budget and playtest time were limits too. Rotating long orchestrator conversations and keeping state outside the conversation helped control context growth. Starting more agents was no longer the obvious way to ship faster.
The next step is scaling verification to match what the fleet can produce.
I built Planestro to standardize the way I prompted agents and remove opportunities for manual errors. Once it was working, it unlocked productivity I hadn’t realized was possible, while preserving the quality standards I expect from my products.
That increased throughput made verification capacity the next engineering problem to solve. Planestro now distinguishes time spent waiting in the CI queue from an actual test failure, and uses queue depth and admission control to manage incoming work. Those changes help the fleet operate within today’s limits. Expanding that capacity while preserving integration tests, review gates, and playtesting is the next stage of the work, and the subject of the third field note.
The first step: change-aware CI.
I have a plan for the next iteration; implementation hasn’t started yet. First, measure runner queue time, suite-lock waiting, and test execution separately, and fix the flaky checks that create avoidable reruns. Then map changes to the checks they can affect.
The selector will start in shadow mode: propose a smaller set of checks while the full suite still runs, so I can measure both missed failures and potential time savings. Targeted selection on epic branches comes only after that trial meets its safety and benefit gates. Unmapped or ambiguous changes fall back to the full suite, with full integration checks retained for gameplay changes reaching develop and main.
Test selection belongs in Hurricane Harrison’s CI workflow. Planestro stays project-neutral, recording validation evidence and making queue and execution times visible. The goal is to spend verification capacity where it matters while preserving the checks that protect the finished product.
The engineering is
in the boundaries.
Hurricane Harrison gave this work a demanding test bed: generated levels, physics timing, art handoffs, and one person accountable for the result. The questions carry into any team putting agents into a delivery pipeline.
Delivery is the outcome.
Hurricane Harrison shipped six App Store releases, from v1.0.0
through v1.5.0. The workflow carries changes from feature branches
through develop and main, uses tags for
versioning, and routes CI to hosted and self-hosted runners.
Planestro grew around that delivery process.
One operator, by design
Planestro currently operates on behalf of one developer, coordinating agents across that developer’s projects. That is a deliberate design choice. Its plans, approvals, credentials, and operations all belong to one person who remains accountable for what the system does.
I’m preparing to open-source it because it has become useful enough in my own work that I want to share it with other engineers who might recognize the same need. They may already have a backlog and a good set of coding tools, but still spend considerable time coordinating the steps between deciding what to build and getting a tested change merged. That was the work my personal playbook captured, and it is the work Planestro now helps me manage.
Its backlog and planning features serve that purpose. They were never intended to replace Jira or Trello. If your team assigns you a Jira ticket, you can create a corresponding item in Planestro, or break it into several dependent items. Planestro helps you plan and carry out that work through implementation, review, CI, integration, and cleanup. Your team’s tracker remains where the shared requirements and priorities live.
That scope matters because Planestro can start agents and coordinate operations that change repositories, create pull requests, merge code, and remove worktrees and branches. Sharing one instance would require explicit rules about whose credentials each operation uses, who can approve which changes, and how one person’s agents are isolated from another person’s work. Switching from subscriptions to API keys would not, by itself, answer those questions.
The current design assumes one trusted operator. Other engineers can run their own instances for their own work, including work they contribute to a team. A shared instance would require a separate design for identity, permissions, and execution isolation. That remains outside the current scope while I focus on making the workflow I use reliable and useful to others.
Why this matters beyond games
For an engineering team, each mechanism answers a practical question about putting agents into production delivery.
| Concern | How Planestro handles it | Why it matters to a team |
|---|---|---|
| Governance | Plan approval starts with me. Delegation requires both the global setting and permission on the item; two fix rounds return the decision to a person. Review holds keep PRs in draft. | Keep people involved where their judgment matters without making them approve every operation. |
| Least privilege | Agents propose orchestration changes through narrow tools. The engine owns pushes, PRs, merges, and state; 27 action tools refuse requests on an item a person has taken over. | Give agents useful work with enforceable limits on their authority. |
| Auditability | Git tracks Markdown state with a semantic commit for each change, such as “HH-143: approve plan”. | Reconstruct the decisions behind a change. |
| Recoverability | Turns and wakes are rebuilt from persisted state, so sessions can be restarted or rotated. | Keep work recoverable across crashes, dropped connections, and session replacement. |
| Safe operations | Cleanup requires a fresh fetch, a merged PR, and ancestry checks on every tip. Dry-run mode and a global pause provide additional controls. | Require positive evidence before destructive automation proceeds. |
| Quality | Behavioral contracts are backed by mutation and negative tests. Repeatable playtest bugs become regression checks. | Evaluate agent-written code against explicit behavior. |
| Cost control | Per-turn cost tracking, usage and disk guards, orchestrator rotation, and branch-based runner routing make resource use visible and controllable. | Understand what throughput costs and where resources are consumed. |
| Capacity | Admission control responds to limits in CI, budget, and review capacity. | Match the rate of new work to the system’s ability to verify and integrate it. |
Planestro Design Principles
Put autonomy behind explicit boundaries.
Keep human gates where product judgment or approval matters. Let deterministic code own shared state, lifecycle transitions, and integration, with agents applying judgment through bounded tools.
Verify behavior, not implementation details.
Use contracts over coordinates and check the effective generated layout. Mutation tests deliberately introduce bad conditions to demonstrate that the checks catch them.
Turn product judgment into lasting requirements.
Playtesting produces rules the fleet can reuse. “Connected, never merely near” means a connection must exist in the generated geometry, not just look close. “Write Help last” keeps documentation tied to mechanics I have played and approved.
Park questions; release processes.
Persist the question, end the turn, and let the process exit. Resume with the answer when it arrives. Human review should not require an agent process to stay alive or a promise to remain pending.
Require proof before destructive cleanup.
A fresh fetch, a merged PR, and ancestry checks establish that the work is preserved before a worktree or branch is deleted. A dirty tree or ambiguous evidence blocks cleanup.
Keep recovery independent of agent memory.
No agent session holds the only copy of anything needed to recover. Every turn and orchestrator wake is rebuilt from persisted state, so sessions can be restarted or replaced. Files are memory. Conversations are cache.
Carry work through to delivery.
Coordinate implementation, review, CI, integration, release, and cleanup. Match new work to verification capacity, and turn incidents into tested improvements that help the next item through the pipeline.
This is the part of agentic engineering I find most interesting: designing the operating model that lets capable agents do useful work, while keeping responsibility for the outcome.