Skip to content

ADR-0105: Autonomous Convergence via Decomposition

  • Status: Proposed
  • Date: 2026-07-12
  • Supersedes: ADR-0091 (Fold Epic Completion Sweep into Epic Monitor) — this decision re-splits epic completion out of the monitor tick that ADR-0091 folded together: completion becomes event-driven (EpicManager._try_auto_close) plus a separate EpicSweeperLoop (ADR-0081), so the monitor tick no longer sweeps.
  • Superseded by: none
  • Amends: ADR-0084 (Auto-Agent as a Universal, Persistent, Root-Cause HITL Gate) — keeps its architecture, interception model, AND its terminal human-required hand-off; inserts an autonomous decomposition step before that terminal, so human-required fires only for the genuinely unmergeable / undecomposable dead-end.
  • Related: ADR-0002 (label state machine); ADR-0050 (auto-agent pre-flight loop); ADR-0029 (caretaker loop pattern); ADR-0044 (recursion safety); ADR-0051 (iterative review convergence); ADR-0059 (advisor / council pattern); ADR-0032 (wiki knowledge base informing the council).
  • Enforced by: tests/test_issue_decomposer.py; tests/test_decomposition_ensemble.py; tests/test_auto_agent_decompose_terminal.py; tests/test_decomposition_depth_cap.py; tests/scenarios/test_decompose_to_converge_scenario.py; tests/sandbox_scenarios/scenarios/s54_decompose_to_converge.py. Nested-lineage follow-up (#9757): tests/regressions/test_epic_lineage_nested_convergence.py; tests/regressions/test_epic_sweeper_lineage_gate.py; tests/regressions/test_epic_manager_lineage_propagation.py; tests/sandbox_scenarios/scenarios/s55_nested_decompose.py.

Note (2026-08-21, flag-rot cleanup): the intake-triage decomposition vector this ADR's Context and Decision reference in present tense (Triage.run_decomposition, TriagePhase._maybe_decompose, and the epic_decompose_on_intake_enabled gate that #11298 had already turned default-OFF) has since been removed outright. The stall-path decomposition this ADR introduces (preflight/decompose_terminal.py + DecompositionCouncil + IssueDecomposer) is now the sole decomposition mechanism, and the §4 depth-counter rationale no longer needs to span two vectors — the only split vector left is the depth-cap-bound stall path. The prose below describes the pre-removal code and stands as written history.

Context

ADR-0084 made the Auto-Agent a universal gate that intercepts hitl-escalation issues, attempts an autonomous fix, and pages a human only for genuinely novel failures. Its terminal, when the auto-agent's own budget (auto_agent_max_attempts, default 3) is spent, is to stamp human-required (auto_agent_preflight_loop.py:_process_one, preflight/decision.py:apply_decision) — a human-owned state.

In practice the gate escalates too often: a change that is merely too broad or ambiguous to converge as one unit hits an attempt cap and pages a human, when it should have been broken into tractable pieces the factory can finish itself. That is routine toil, not a real request for help.

The operating requirement: eliminate routine human escalation, and reserve human-required for the genuinely unmergeable / undecomposable dead-end — the case where the factory truly is stuck and a human should step in (a real bug it can't fix, a missing external dependency, ambiguous intent no split resolves, a broken factory). HITL is correct there; it is wrong for a change that only needed splitting.

The lever HydraFlow already has for "too big to converge as one unit" is decomposition: Triage.run_decomposition already splits an issue into an epic + 2–6 child specs, and EpicCompletionChecker/EpicSweeperLoop already auto-close a parent when its children merge. Today that path fires only on intake complexity, never on retry exhaustion — so a stall that a split would fix escalates to a human instead.

Decision

When a change exhausts the auto-agent's budget, try to decompose it into smaller child issues before paging a human. If it splits, the parent converges when its children converge (no human). If it cannot split, it falls through to ADR-0084's existing human-required terminal — now reached only for the genuine dead-end. This inserts a step; the interception model, retry semantics, and playbooks are unchanged.

  1. Insert decompose before the terminal. At the two human-required sites, first attempt decomposition. human-required is applied only when decomposition is declined or exhausted — so it survives as the honest "factory is stuck, needs a human" signal, no longer fired for merely-too-big changes.

  2. Council-based, doc-informed decomposition. The split is decided by a council, not a single LLM shot — reusing the established council / advisor pattern (cf. the ADR-review council in adr_reviewer.py, ExpertCouncil, ADR-0059). Two passes:

  3. Direction — propose candidate slicings from distinct lenses (architectural / layer boundaries; isolate-the-failing-part; vertical independently-shippable slices), given the stall context (failing stage, blocked_reason, diagnosis, diff-so-far, review findings).
  4. Validation — judge the chosen split: are the children independently-shippable, non-overlapping, genuinely more tractable than the parent (not clones), and consistent with the architecture? If no sound split survives, the council declines (should_decompose = false) → the HITL floor.

The council is informed by the docs: the relevant ADRs (resolved from the change's touched files via the ADR cross-reference) and the relevant wiki entries (patterns / gotchas / testing for the affected area, per ADR-0032) are fed into its context, so the slicing respects accepted architectural decisions and reuses accumulated knowledge instead of re-deriving it. Extract the create/link/register plumbing from TriagePhase._maybe_decompose into a standalone IssueDecomposer the council drives; the existing single-shot Triage.run_decomposition stays the intake path.

  1. Reuse the epic machinery — and carry the learning forward. On should_decompose = true: create the epic + children (children enter at hydraflow-find and run the full pipeline), register_epic(auto_decomposed=True), then close the stuck issue as decomposed, close its superseded PR, and let workspace_gc_loop reap its worktree — the exact terminal pattern intake decomposition already uses. Parent rollup is free for the depth-1 case (EpicCompletionChecker + EpicSweeperLoop close a parent when its children close). Nested cascade (a depth-2 parent closing only once its grandchildren merge) was not free at P1: EpicState had no epic-to-epic lineage, so a re-decomposed child was marked decomposed/closed immediately and the root could close before its grandchildren finish. P1 shipped max_decomposition_depth default 1 to sidestep that. The follow-up (#9757) has since landed: EpicState.parent_epic/superseded_issue lineage + the sweeper gate + EpicManager._propagate_epic_close make nested cascade correct, so the default is now 2 (see Consequences).

Two things the naive "split and discard" version gets wrong, so they are part of the decision: - Carry the failure signal down. The parent stalled for a reason (the stall context). Each child body embeds that signal — "the parent failed on X; this slice must not repeat Y" — so children don't re-derive the parent's mistake and re-stall. Without this, decomposition just produces smaller clones. - Salvage partial work. A stall at implement/review usually has a partial diff, some of it good. The council sees the diff-so-far and, when part of it is sound, scopes a child as "land the working slice" rather than throwing all progress away and paying to rebuild it. - Idempotent. The auto-agent runs once per tick and can be interrupted mid-decompose. The terminal checks for an already-registered auto-decomposed epic for this issue before creating one, so a retry never duplicates the epic/children.

  1. Bounded depth AND fan-out (moderate budget). Depth alone is not enough: with up to 6 children per split and depth 2, the naive worst case is 6×6 = 36 leaf pipelines from one root. Two bounds:
  2. max_decomposition_depth (default 2; the epic-lineage follow-up #9757 landed, so nested cascade is correct) — a decomposed child that also stalls may split once more, then hits the floor. Depth is tracked per epic/child; le=5 bounds the chain.
  3. max_total_decomposition_children per root (default ~8) — a blast-radius cap across the whole tree, so fan-out can't explode regardless of depth. Hitting it → HITL floor for the remainder.

The depth counter must span both decomposition vectors. Auto-generated children re-enter at hydraflow-find and run intake triage, which has its own complexity-gated decomposition. An auto-child is stamped so intake decomposition does not re-split it uncounted — every split, stall-path or intake-path, increments the same depth. The pre-decompose retry caps (implement 3 / review 2 / auto-agent 3) are unchanged — the "fair shot" before a split.

  1. The floor: HITL, for the genuine dead-end. When should_decompose = false (not decomposable) OR depth is exhausted, fall through to ADR-0084's human-required terminal + a full diagnostic comment. This is the correct use of HITL — the factory has exhausted its autonomous options (retries, auto-agent, decomposition) on a change that is neither mergeable nor splittable, so a human genuinely needs to step in. The change from today is selectivity, not removal: HITL becomes rare and meaningful instead of routine.

Consequences

  • HITL becomes rare and meaningful. A pipeline change reaches merged, decomposed→merged, or — only when neither is possible — human-required with a diagnostic. The routine "too-big → human" escalations are gone; the ones that remain are genuine "factory is stuck, needs help" signals, which is exactly what HITL is for.
  • Correct use of the human. Decomposition turns too-big / ambiguous into tractable; it cannot manufacture a merge for a genuinely-impossible change (real bug it can't fix, missing external dependency, unresolvable ambiguity). Those should reach a human — that is the case HITL exists for. The design narrows HITL to that case, it does not abolish it.
  • The HITL that remains is tiny and well-scoped. This is the real prize, beyond frequency: when a change still can't converge, the human is handed the one small failing slice (a child scoped to the exact stuck part) with a diagnostic — not a giant vague parent. Decomposition isolates the irreducible difficulty to its minimum, so the human's job is small, specific, and fast even in the residual case.
  • Cost shifts from human-time to compute (double-bounded). A decomposed issue spawns child pipelines — more $ than a ~zero-compute human page. Bounded by BOTH max_decomposition_depth and max_total_decomposition_children per root, so fan-out cannot explode; exceeding either → HITL, not more burning. A council run (a few LLM calls) per stall is itself a cost, justified because a sound split avoids the far larger cost of children that re-stall.
  • Recursion must be capped on depth AND fan-out. Neither exists today; both are load-bearing. The depth counter spans the stall path and the intake-triage decomposition path, or an auto-child could be re-split uncounted.
  • Scope is phased. P1 inserts decompose before the auto-agent's terminal (main pipeline: plan/implement/review). P2 routes the ~7 side loops that file hitl-escalation directly (discover, contract-refresh, corpus-learning, …) through the same decompose-first terminal, so routine escalation is eliminated everywhere. (No parked-revisit phase — the floor is the existing HITL terminal.)
  • State reconciliation. A stuck issue owns a branch/worktree/PR + attempt counters; on decompose, the terminal closes/cleans all of it, reusing the intake path's proven close-and-supersede handling.
  • Default max_decomposition_depth is 2; nested cascade converges (was P1-deferred, #9757 landed). A review found that nested (depth ≥ 2) decomposition could close a root epic before its grandchildren finished: EpicState had no epic-to-epic lineage, so a re-decomposed depth-2 child's completion was invisible to the root's convergence check. P1 shipped default 1 to sidestep that. The follow-up (#9757) adds EpicState.parent_epic/superseded_issue: the epic sweeper (EpicSweeperLoop._try_sweep_epic) refuses to treat a decomposed-closed child as resolved until its replacement epic's GitHub issue closes, and EpicManager._propagate_epic_close cascades a replacement epic's close up the parent_epic chain (N-level) so a superseded child only counts as resolved once its replacement epic finishes. With that in place the default rose to 2 (le=5 still bounds the chain).