ADR-0098: Convergence oscillation caretaker (Phase 2d)¶
- Status: Accepted
- Date: 2026-07-01
- Refines: ADR-0094 (two-level convergence: Gate + ConvergenceLedger)
- Extends: ADR-0096 (boundary verdict recording)
- Enforcement: enforced
- Enforced by: pytest:tests/scenarios/test_convergence_oscillation_mockworld.py
Precedent: Limit-cycle / hunting detection in control systems — recognizing sustained oscillation about a set-point rather than convergence to it Divergence: classical hunting detection watches one controlled variable's trajectory, but here the signal is a repeated cross-boundary finding-signature set spanning Triage/Shape/Plan/Review laps that no single stage's own detector can see, so a read-only caretaker consumes the pipeline-wide verdict history and escalates once to HITL (receipt: ADR-0096, #9706)
Context¶
Phase 2b (ADR-0096) made the ledger record LOOP_BACK/ADVANCE verdicts at each pipeline boundary (Triage, Shape, Plan, Review), giving the ledger a pipeline-wide verdict history for the first time. That history is observability data: no existing consumer acted on cross-boundary oscillation patterns.
The review phase (ADR-0094/0095) escalates to HITL when the outer lap budget (max_convergence_laps) is exhausted or when detect_outer_oscillation detects repeated finding-signature sets across review laps. Both signals are anchored at the Review boundary. Two classes of stuck issue are not caught:
- Issues that ping-pong across Triage, Shape, and Plan without reaching the review lap cap. Each stage routes them back, but no global observer detects the pattern.
- Issues that reach Review, trigger oscillation, but have not yet closed enough laps to exhaust the lap cap. The cap and oscillation detector are review-lap-anchored; a cross-boundary pattern that spans fewer review laps than the cap is invisible to them.
Phase 2d adds ConvergenceOscillationLoop, a background caretaker that consumes the cross-boundary verdict history accumulated by Phase 2b and escalates stuck issues to HITL before they exhaust the lap cap or loop forever in the pre-review stages.
Decision¶
Detection: ConvergenceLedger.detect_cross_boundary_oscillation¶
The detection method fires on EITHER of two complementary signals:
Temporal signal (post-review oscillation): detect_outer_oscillation(window) returns True when the last window review laps produced identical, non-empty finding-signature sets. lap_signatures unions all boundary findings at mark_lap, so this signal is cross-boundary-aware for issues that reach review. The caretaker gates this arm via detect_cross_boundary_oscillation(include_temporal=...), firing it only when laps == 0 (pre-review, where the snapshot signal has no review data) or laps >= max_convergence_laps (review budget exhausted). Between those bounds (0 < laps < max_convergence_laps) the temporal arm defers to the review loop's own detect_outer_oscillation, so the caretaker and the review loop do not race on live laps (PR #9706).
Snapshot signal (pre-review churn): at least min_loopback_stages distinct stages among {triage, shape, plan} currently have last_verdict == "LOOP_BACK". This catches cross-boundary churn in issues that have not yet closed a review lap, where the temporal signal has no data.
Either signal alone is sufficient to flag an issue as oscillating. Default values: window=2, min_loopback_stages=2.
The loop: ConvergenceOscillationLoop¶
The loop (src/convergence_oscillation_loop.py, class ConvergenceOscillationLoop(BaseBackgroundLoop), worker name convergence_oscillation) runs on a configurable interval and performs three steps:
- Enumerate all per-issue ledgers via
StateTracker.iter_convergence_ledgers(), a new public accessor added tosrc/state/_convergence.py. - Skip any ledger where
ledger.convergedis True orledger.oscillation_escalatedis True. - For each remaining ledger, call
detect_cross_boundary_oscillation. If it fires, create a companion HITL issue labeled with the HITL escalation label andconvergence-oscillation, then setledger.oscillation_escalated = Trueand persist.
The dedup flag (oscillation_escalated) is set only AFTER a successful create_issue call, so a failed create retries on the next interval. A failed create that raises a credit or bug exception is not swallowed: reraise_on_credit_or_bug(exc) is called in the broad except block before any fallback, propagating CreditExhaustedError and likely-bug exceptions upward (per the dark-factory contract in docs/wiki/dark-factory.md section 2.2).
The loop makes no LLM calls (LONG_LLM_CYCLE = False). loop_fitness is HOUSEKEEPING / INSUFFICIENT_DATA: the caretaker has no clean per-item acceptance signal because escalation is a one-shot side effect, and the absence of oscillation is the healthy state.
Read-only-except-dedup contract¶
The loop NEVER calls mark_lap, record_gate_result, increment_attempts, or recompute_converged. Those methods remain exclusively owned by the review phase. The loop's only ledger write is the oscillation_escalated dedup flag. This contract keeps boundary phases simple: each boundary records its own verdict and drives its own lap accounting; the caretaker is a read-only observer that fires a one-shot side effect when a global pattern is detected.
Control and safety¶
Two-layer kill switch per ADR-0049: an in-body _enabled_cb callback checks the live config value, and the config field convergence_oscillation_loop_enabled (default True) provides operator-level control. The loop also checks dry_run before creating any issue: in dry-run mode, detection runs but no issues are filed.
Configuration¶
All fields are env-overridable:
convergence_oscillation_interval(default 3600, ge=300): polling interval in seconds.convergence_oscillation_loop_enabled(default True): operator kill switch.convergence_oscillation_window(default 2): number of review laps compared by the temporal signal.convergence_oscillation_min_loopback_stages(default 2): minimum distinct stages at LOOP_BACK required for the snapshot signal.
Relationship to review escalation¶
The caretaker complements, and does not replace, the review phase's lap-cap escalation. The lap cap catches issues that exhaust their review budget regardless of oscillation pattern. The caretaker catches two cases the lap cap misses: (a) cross-boundary churn that never reaches the review lap cap (snapshot signal, laps == 0), and (b) review-lap oscillation once the lap budget is exhausted (laps >= max_convergence_laps). While 0 < laps < max_convergence_laps, case (b) is left to the review loop's own detect_outer_oscillation on live laps; the caretaker's temporal arm re-engages only at the cap (PR #9706), so the two never double-escalate the same live lap. Both mechanisms can trigger on the same issue without conflict: oscillation_escalated prevents duplicate caretaker escalations, while the lap cap follows its own logic independently.
Rules and consequences¶
- Read-only contract is absolute. The loop must not call any ledger method that advances lap state or alters verdicts. The only permitted ledger write is
oscillation_escalated. Violating this would corrupt the review phase's lap accounting. - Dedup flag set after, not before, the create. A pre-set flag would suppress retries on a failed create, leaving the issue unescalated with no way to recover without operator intervention.
- Credit and bug exceptions propagate. The broad
exceptblock must callreraise_on_credit_or_bug(exc)before any fallback. SwallowingCreditExhaustedErrorsilently burns attempt budget against an exhausted billing signal. - Dry-run is respected. The loop may detect oscillation in dry-run mode, but it must not create any issue. Dry-run logs the detection for observability without side effects.
- Config knobs are additive, not subtractive. Tightening
convergence_oscillation_windoworconvergence_oscillation_min_loopback_stagescauses earlier escalation. Loosening them reduces sensitivity. Neither affects the review lap cap. oscillation_escalatedis terminal. Once set, the caretaker ignores that ledger. Clearing it to re-escalate requires operator intervention on the state file.
Scope (Phase 2d)¶
This ADR covers:
- The
detect_cross_boundary_oscillationdetection method and theoscillation_escalateddedup flag, both onConvergenceLedger(src/models.py). - The
iter_convergence_ledgerspublic accessor andmark_oscillation_escalatedhelper onStateTracker(src/state/_convergence.py). ConvergenceOscillationLoopand its full orchestrator and registry wiring (src/convergence_oscillation_loop.py).- The full test pyramid: unit tests for detection and the loop, a MockWorld scenario, and a sandbox e2e scenario.
This completes the Phase 2 arc. Phase 2a gated the approve path. Phase 2b added cross-boundary verdict recording. Phase 2c migrated attempt counters into the ledger. Phase 2d adds the oscillation caretaker that consumes Phase 2b's verdict history. Nothing in Phase 2d is out of scope as deferred: all invariants (functional_areas.yml, scenario catalog, event reducer, orchestrator wiring, registry, AST ratchets) are addressed in Tasks 1-5.
Alternatives considered¶
-
Observability only, no auto-escalation. Emit a metric or log when oscillation is detected, and leave escalation to a human watching a dashboard. Rejected. The whole point of the caretaker is an autonomous safety net. An operator watching a dashboard reintroduces the human-in-the-loop latency the pipeline exists to eliminate. Issues stuck in oscillation can consume resources indefinitely without an automated response.
-
Reuse the review-lap temporal detection only, no snapshot signal. Run
detect_outer_oscillationacross all ledgers and skip the snapshot check. Rejected. This misses the pre-review cross-boundary churn case: issues that loop between Triage, Shape, and Plan without ever closing a review lap have no lap data for the temporal signal to operate on. Those issues would loop forever under this approach. -
Let each boundary drive escalation inline when it detects repeated LOOP_BACK. Each phase checks its own LOOP_BACK count and escalates when it exceeds a threshold. Rejected. This duplicates escalation logic across three boundary phases, makes the threshold a per-phase config that must be kept consistent, and produces escalation events from multiple sites for the same issue. A centralized read-only caretaker sees the full cross-boundary picture, escalates once, and keeps boundary phases responsible only for their own verdicts.
When to supersede this ADR¶
Supersede when: the detection algorithm changes (for example, adding a frequency-based signal or making the window adaptive); the oscillation_escalated dedup strategy changes (for example, to allow re-escalation after a configurable cooldown); the loop gains LLM calls or a different fitness classification; or Phase 3 folds oscillation handling into a unified convergence runner that replaces the caretaker pattern.
Source-file citations¶
src/models.py:ConvergenceLedger.detect_cross_boundary_oscillation,ConvergenceLedger.oscillation_escalated.src/state/_convergence.py:StateTracker.iter_convergence_ledgers,StateTracker.mark_oscillation_escalated.src/convergence_oscillation_loop.py:ConvergenceOscillationLoop, workerconvergence_oscillation.tests/scenarios/test_convergence_oscillation_mockworld.py: MockWorld scenario exercising the full caretaker flow.tests/sandbox_scenarios/scenarios/s51_convergence_oscillation.py: sandbox e2e scenario.