Skip to content

Dark-factory engineering

Doctrinal page — hand-authored, not LLM-maintained. Unlike the other files in docs/wiki/, this page intentionally has no json:entry metadata blocks. It is curated prose distilling load-bearing operating doctrine, not auto-ingested per-issue insights. RepoWikiLoop does not rewrite this file; the wiki injector still reads it into runner prompts as architectural context.

The lights-off operating contract: any HydraFlow-managed project meeting the spec runs autonomously, with humans paged only for raging fires. This entry distills the load-bearing conventions that make that contract real, the recurring footguns that break it, and the pattern for delivering substantial features that survive production.

§1 — The contract

Auto-Agent (ADR-0050) is the issue-queue layer of the contract: every hitl-escalation issue is intercepted by AutoAgentPreflightLoop, which attempts autonomous resolution before the issue surfaces to a human via human-required. The trust fleet (ADR-0045) is the runtime layer: ten caretaker loops watch for drift, anomalies, principles violations, RC budget overruns, etc., and either auto-repair or escalate. ADR-0049 is the universal kill-switch convention that lets operators flip any loop off live.

What "dark-factory ready" means concretely:

  • Every escalation has an autonomous fix-attempt path before a human sees it.
  • Every loop is independently kill-switchable at the UI without a redeploy.
  • Every cost is observable on the dashboard with attribution to a source.
  • Every repair is auditable post-hoc via JSONL streams.
  • Every "new thing" inherits or replicates the load-bearing conventions — there is no single load-bearing convention you can skip without breaking the contract.

§2 — Load-bearing conventions for new code

2.1 New caretaker loop checklist

Every new loop must:

  1. Five-checkpoint wire. config.py field + env override, service_registry.py import + dataclass field + construction + ServiceRegistry kwarg, orchestrator.py bg_loop_registry + loop_factories, src/ui/src/constants.js EDITABLE_INTERVAL_WORKERS + SYSTEM_WORKER_INTERVALS + BACKGROUND_WORKERS, src/dashboard_routes/_common.py::_INTERVAL_BOUNDS, tests/scenarios/catalog/loop_registrations.py builder + entry. Verify with tests/test_loop_wiring_completeness.py (regex auto-discovery). Two more gates fire only in FULL make quality (learned building GateHealthLoop #9974 — each cost a 15-minute bake to discover): an explicit loop_fitness override (tests/test_loop_fitness_completeness.py — SCORED via helper or HOUSEKEEPING per ADR-0093; never add to _GRANDFATHERED; mirror adr_conformance_loop.py), and a services.<name>_loop = FakeBackgroundLoop() stub in tests/orchestrator_integration_utils.py::build_scripted_services (the orchestrator's bg_loop_registry dereferences every registered loop, so four integration tests AttributeError without it).

  2. ADR-0049 in-body kill-switch gate at the top of _do_work:

    if not self._enabled_cb(self._worker_name):
        return {"status": "disabled"}
    
    This is universal — no exceptions. The 18-loop retrofit (PR #8430) was this convention being applied codebase-wide.

  3. Static config gate (*_enabled env var) for deploy-time disable that doesn't require the UI being up. Defense-in-depth alongside the in-body enabled_cb gate.

  4. Functional area assignment in docs/arch/functional_areas.yml. The architecture tests (tests/architecture/test_functional_area_coverage.py) fail if a loop or port is unassigned.

  5. Architecture-generated docs regen (make arch-regen). The tests/architecture/test_curated_drift.py test fails on stale generated docs — easy to forget after adding a new module.

2.2 Subprocess runner conventions

Either inherit from BaseRunner (you get auth-retry + telemetry + tracing context for free) OR replicate explicitly:

  • 3-attempt auth-retry loop with exponential backoff (5s, 10s, 20s) on AuthenticationRetryError from runner_utils. Transient OAuth blips shouldn't burn the per-issue attempt cap.
  • reraise_on_credit_or_bug(exc) in the broad except to propagate CreditExhaustedError and terminal AuthenticationError from subprocess_util. Without this, the loop continues ticking against an exhausted billing signal.
  • PromptTelemetry.record(source=...) on every attempt for cost rollup attribution. Use a unique source string so the dashboard can break out spend per runner.
  • Never-raises contract: every failure path returns a typed result (e.g., PreflightSpawn(crashed=True, ...)), never propagates a generic RuntimeError. The caretaker loop's outer handler shouldn't need to know about subprocess internals. This bullet does NOT outrank the one above it. "Never raises" means never raises an ordinary failure; infra-fatal errors still propagate. Read literally it says the opposite, and that reading is what produced the #11666 sweep's largest finding: TranscriptSummarizer.summarize_and_comment documented itself as "Never raises — all errors are logged and swallowed" and duly buried CreditExhaustedError from its own summarization spawn as return False. A degraded result is only honest when the next call could succeed; after credit exhaustion it cannot.
  • Fix the whole chain, not the innermost site. In that same finding, all five callers of the summarizer re-swallowed on top of it — two via except (RuntimeError, OSError), which catches CreditExhaustedError because it subclasses RuntimeError. Adding the guard only at the inner site would have moved the swallow up one frame and looked fixed. Walk the callers.

The Auto-Agent partial-landing → wiring follow-up (PRs #8431 → #8439) exposed all four of these as load-bearing — the first runner cut missed auth-retry AND reraise_on_credit_or_bug, and both were caught only by fresh-eyes review.

The gate: tests/test_loop_credit_reraise_completeness.py now ratchets this over every module in src/, not just src/*_loop.py, and its call graph resolves across a decomposed loop's sibling modules rather than one file at a time (#11666). Both grandfather lists are empty and must stay that way. Before #11664 the guard was believed present everywhere while the #6855 site had none — its regression test was anchored to a line window that had drifted off the method, so it passed vacuously for months. Anchor structural tests on SYMBOLS, never on line numbers.

2.3 Audit-on-everything

JSONL audit stream is the source of truth. Use file_util.append_jsonl (does fsync) + file_util.file_lock (advisory lock) for durability. StateData fields cache for fast dashboard reads; the JSONL is canonical. This is the spec §6.3 contract for PreflightAuditStore — and the same pattern applies to any new audit/cost/event JSONL stream.

2.4 Observability-first guardrails

Wire caps (cost, wall-clock, daily budget) into code paths but default to None (unlimited). Dashboard surfaces data without alerting. Operator decides when to impose policy. From spec §5.1 of the Auto-Agent design: "observability-first; operator can set when needed". Avoid premature gating that wastes operator attention on cap-hits before they have data to know what the right cap is.

2.5 Honor-system + post-hoc CI enforcement

When runtime can't enforce a rule (e.g., file-path restrictions on a Claude Code subprocess where the CLI flag operates on tool names not paths), document clearly in the prompt envelope AND rely on principles-audit + CI to catch violations. Don't lie about enforcement boundaries — operators reading docs during incident triage need to know what's runtime-enforced vs what's post-hoc-audited. The Auto-Agent _envelope.md revision (src/hydraflow_resources/prompts/auto_agent/_envelope.md) is the reference: clearly separates "Enforced by the Claude Code CLI" from "Enforced post-hoc by CI / principles audit".

2.6 Partial-landing visibility

If you ship scaffolding with a placeholder for a load-bearing piece, make the placeholder OPERATIONALLY OBSERVABLE: zero spend on the dashboard, zero resolution rate, distinct status string in the loop's payload. Document in the ADR's Consequences section, not just a code TODO. Auto-Agent shipped with a placeholder _build_spawn_fn in PR #8431; the dashboard's resolution_rate=0 + spend_usd=$0 was the operational signal. PR #8439 removed the placeholder, and ADR-0050 §Consequences was updated to mark the wiring landed.

2.7 Sub-label deny-list for recursion safety

Caretaker agents that act on the system shouldn't act on the system that judges them. Auto-Agent's deny-list (auto_agent_skip_sublabels = ["principles-stuck", "cultural-check"]) prevents auto-agent attempts on principles violations — letting auto-agent "fix" a principles audit failure by editing the auditor would defeat the audit. Hard tool restrictions in the prompt envelope reinforce this for file-level rules (auto_agent_preflight_loop.py, principles_audit_loop.py, ADR-0044/0049/0050 implementation files).

§3 — The production-readiness convergence loop

For substantial features (new loop, new runner, spec → implementation):

  1. Brainstorming → spec → plan → implementation. Standard workflow.
  2. Per-task review during implementation. Subagent-driven development (superpowers:subagent-driven-development) does spec-compliance review
  3. code-quality review per task. The trust-arch and auto-agent features used this — every task got 2 reviews.
  4. Fresh-eyes review iterations after implementation. A reviewer who doesn't see the conversation context catches things you've grown blind to. Plan for 2–3 iterations before merge. Each pass finds fewer issues. Convergence = next pass finds nothing material.
  5. Smoke-test before merge. make quality is the actual gate, not pytest tests/test_*.py. Architecture tests (test_functional_area_coverage, test_curated_drift, test_loop_wiring_completeness, test_port_conformance) catch a class of issues unit tests don't.
  6. PR-merge collisions. When main moves while your PR waits: git rebase origin/main -X theirs for arch-generated conflicts, then make arch-regen, then re-CI. Don't try to manually merge generated files — they're stale baselines, not real conflicts.

Feature-by-feature data points: - Trust-fleet (PR #8390) → 5 audit passes to convergence. - Auto-Agent spec (PR #8431) → 3 spec review + fix iterations. - Auto-Agent subprocess wiring (PR #8439) → 3 fresh-eyes review iterations.

The convergence point is reliably ~3 passes for substantial work. Plan for it; don't merge before it.

Sandbox-tier expectations (added 2026-04-28 — ADR-0052)

For substantial features, the convergence loop now extends to the sandbox tier:

  • Every runnable sandbox scenario must pass on the rc/* promotion PR before the staging→main merge can complete. CI gates this via the sandbox-full job. Placeholder scenarios are removed from the runnable catalog until they can assert real behavior.
  • Failures auto-dispatch SandboxFailureFixerLoop, which gives the auto-agent up to 3 attempts before escalating to the System tab HITL queue (via /api/sandbox-hitl).
  • Nightly sandbox runs catch slow drift; failures open hydraflow-find issues per the 3-strikes-then-bug pattern.

The same MockWorld substrate (src/mockworld/fakes/) backs both in-process Tier 1 and sandbox Tier 2; Port↔Fake conformance tests keep them aligned.

§4 — Recurring footguns

4.1 Subagent claims DONE without committing

Subagents sometimes report DONE with edits applied but not committed, or with a partial commit that left some files staged. Always run git status --porcelain and git log -1 --stat after a subagent reports DONE before moving on. Hit twice during the auto-agent work (T12 wiring and T13 close-reconciliation tasks).

4.2 AsyncMock hides PRPort method-name typos

Tests that use AsyncMock(pr) auto-create any attribute name on access, so a typo like pr.remove_labels(...) (plural) when the real method is remove_label (singular) passes the test but crashes in production. The tests/test_ports.py structural conformance suite is the safety net (it checks each adapter against its Port protocol via runtime_checkable isinstance AND inspect.signature comparison; per-Fake conformance lives in tests/scenarios/fakes/test_fake_*.py since the Fakes moved to src/mockworld/fakes/) — make sure any new method on a real Port is also added to the corresponding Fake AND the conformance suite runs. The C2/C3 critical findings on PR #8439 were exactly this class of break.

4.3 Pyright IDE noise on Pydantic dynamic attrs

Pyright's static analysis can't follow the indirection from self._data: StateData to fields on dynamically-composed mixin classes. Diagnostics like Cannot access attribute "auto_agent_attempts" for class "StateData" are expected noise and tolerated by the build's pyright config. Trust the build, not the IDE diagnostics. Every existing mixin (_flake_tracker.py, _contract_refresh.py, etc.) shows identical IDE warnings while passing CI.

4.4 Ruff strips unused imports during TDD

If you add an import (e.g., from x import field) before the code that uses it, ruff's auto-fix on save strips the import as unused. The fix: append the implementation that uses the import FIRST, then add the import. Or use locally-scoped imports inside test functions when ruff keeps stripping. Already in user memory; surfaces every few tasks.

4.5 Generated-file rebase pain

Conflicts in docs/arch/generated/, docs/arch/.meta.json, etc. on rebase aren't real conflicts — they're stale baselines that need regeneration. Recipe:

git rebase origin/main -X theirs
make arch-regen
git add -A && git commit -m "chore(arch): regen after rebase"
git push --force-with-lease
Hit twice during auto-agent work; both times the -X theirs + make arch-regen recipe resolved cleanly.

4.6 Tests job timing race on auto-merge

The CI Tests job runs the full ~11k-test suite (~7 min). Force-pushes during this window invalidate the run and trigger a fresh CI cycle — which is fine, but gh pr merge will reject the merge as "Pull Request has merge conflicts" if main moved during CI. The --auto flag is the ideal recipe but only works if the repo enables it; otherwise, monitor CI completion and manually merge.

4.7 Cost-watcher operator-override windows

CostBudgetWatcherLoop takes authorship of two reversible actions — the hard-cap kill (set_enabled False) and the soft-band throttle (interval stretch, cost_throttled_workers priors). In both cases an operator change made inside the window is silently superseded on recovery: a worker manually disabled after we killed it gets re-enabled; an interval changed mid-throttle is overwritten by the pre-throttle value (None → cleared to the loop default). Detecting either would need an event log keyed on (name, source, timestamp) for every control-plane write. Accepted trade-off: the windows are short (band crossings move at rolling-24h speed) and recovery restores the pre-window operator intent.

4.8 Pre-push self-check — six recurring avoidable CI reds

Six patterns pass make quality-lite locally but go red on the first CI round, each costing a heal round-trip. All are already CI-guarded — the win is pre-empting them at build time. The implementer prompt now carries these as a "Pre-push self-check" section (src/agent/_runner.py:_SELF_CHECK_CHECKLIST); the same list, with the issue references that surfaced each, lives in gotchas.md ("Pre-push self-check — six recurring avoidable CI reds").

  1. fix( commit → tests/regressions/ delta (P10.6). No regressions test on a fix( subject WARNs → red. Add one, or a Skip-Regression: trailer for a pure refactor.
  2. New run_subprocess*/stream_claude_process call site → sandbox seam. Declare it in mockworld.sandbox_main.SANDBOX_SEAMS or route through an injected-fake seam, or the seam-completeness ratchet goes red.
  3. ADR Enforced by: with multiple checks → bullet lines. **Enforced by:** then - pytest:a / - pytest:b; never inline pytest:a, pytest:b.
  4. Code extraction relocates a # noqa → disturbance ratchet red. Narrow the except to concrete types or hoist the import to module top so no suppression is needed; never bump the baseline.
  5. Moving a cited file → update its ADR Enforced by: citation in the same commit, or ADR-conformance goes red.
  6. Relocating a symbol → repoint tests that patch() it. Grep for patch("oldmodule.symbol") before the move, or the test errors at collection.

4.9 Token-drift baseline regeneration

token_drift.py (#11441) pins a per-source token-share + median-tokens-per-issue baseline and compares each new trailing ISO week against it. GET /api/diagnostics/token-report's drift block degrades to no_baseline (never pinned or the ledger is unreadable/corrupt), insufficient_data (fewer than MIN_BASELINE_WINDOWS — 8, mirroring VitalsThresholds.min_baseline_windows — pinned windows, no issues in the trailing week, or a trailing week the loader could not cover completely — the telemetry read died mid-stream, or audit_retention_days_inference_telemetry falls inside the window so its head may already be pruned; #11581), or stale (pinned more than MAX_BASELINE_AGE — 90 days — ago) instead of ever fabricating a verdict. All three mean the instrument is not watching — drift is not being checked, not that drift was checked and found clean.

Telemetry is loaded by window, never by row count (token_drift. load_window_rows streams inferences.jsonl via PromptTelemetry. iter_inferences and keeps every row of trailing_complete_weeks(now, n)). The original DRIFT_LOAD_LIMIT = 5000 tail cap went blind in exactly the weeks the sensor exists for: 2026-W25 alone carried 24,434 rows, so the open week pushed the trailing week out of the cap within ~1.5 days (insufficient_data, nothing filed) and a cap landing mid-week mis-sampled the shares into a confident wrong verdict (#11581).

Re-pin whenever:

  • No baseline has ever been pinned (no_baseline on the diagnostics panel).
  • The baseline goes stale (>90 days old).
  • A deliberate token-efficiency lever landed (session continuation, cache-aware prefixes, size tiering, …) and the new steady state should become the comparison point, not a drift.
python scripts/pin_token_baseline.py --reason "why you're (re)pinning now"

That's a dry run — it prints the windows, sources, and median series it would pin without writing anything. Add --apply to actually write the ledger (there is no separate --dry-run flag; omitting --apply is the dry run — mirrors scripts/cleanup_phantom_cost.py):

python scripts/pin_token_baseline.py --reason "why you're (re)pinning now" --apply

--reason is required — re-pinning silently discards the ability to detect drift against the old steady state, so state why up front, the same discipline scripts/regen_concentration_baseline.py enforces for the concentration baseline. The ledger lands at <data_root>/calibration/token_baseline.jsonl (append-only, finder_calibration.CALIBRATION_SUBDIR; last row wins), needs at least 8 complete ISO weeks of inferences.jsonl telemetry to pin from (--windows overrides the trailing-week count), and is read back by token_drift.load_and_check_drift on every /token-report request and by ErosionMetricsLoop on each tick that saw new commits, which files one hydraflow-find issue per drifting source per ISO week (#11442, erosion.token_drift_filing); the engine itself stays read-only — no prompt pruning, no config change. Verify a pin landed via curl localhost:<port>/api/diagnostics/token-report | jq .drift.

§5 — Verifying the contract is honored

Auto-discovery tests that fail when a load-bearing convention is broken:

Test What it catches
tests/test_loop_wiring_completeness.py Loop missing one of the five checkpoints
tests/architecture/test_functional_area_coverage.py New loop or port unassigned in functional_areas.yml
tests/architecture/test_curated_drift.py Generated docs out of sync after a source-file change
tests/test_ports.py Adapter/Fake drifts from its Port protocol signature
tests/test_loop_kill_switch_completeness.py Loop without ADR-0049 in-body gate
tests/test_config_consistency.py *_interval config field without matching _INTERVAL_BOUNDS entry

Before marking work complete: run make quality. It runs all of these plus the full suite. Unit tests passing is necessary but not sufficient.

§6 — The meta-pattern

Across every Critical finding caught in review across the last six PRs, the pattern was: a load-bearing convention was something a careful engineer remembers, not something the codebase forces.

The meta-improvement is moving conventions from "remembered" to "structurally enforced" — base classes that auto-apply patterns, scaffold scripts that generate boilerplate with all the conventions correct, conformance tests that catch contract drift, pre-commit checks that block the most common omissions.

See ADR-0051 for the formal "iterative production-readiness review" process. The infrastructure improvements this section originally planned have since SHIPPED — check for the current state before rebuilding any of them (two 2026-07 sessions nearly did):

  • BaseSubprocessRunner (src/runners/base_subprocess_runner.py, #8446) — auto-applies reraise_on_credit_or_bug + telemetry ordering.
  • scripts/scaffold_loop.py + scripts/scaffold_templates/ (#8448) — atomic five-checkpoint patcher; template kept current with the ratchets (kill-switch gate order, loop_fitness override).
  • Conformance ratchets: tests/test_loop_wiring_completeness.py, tests/test_loop_fitness_completeness.py, tests/test_loop_kill_switch_completeness.py, Port↔Fake signature conformance (#8446).

Still open: subagent-verify wrapper, pre-commit arch-regen.

Exception sensor — standing Bugsink up (ADR-0146)

There is no single make start. The sensor is separate infrastructure on purpose — HydraFlow runs perfectly well without it, because an absent DSN gets the no-op adapter — so it is its own command:

Command Starts Port
make run dashboard + UI 5555 (loopback)
make factory the factory loops —
make bugsink-up Bugsink + its database (local only) 8000 (loopback)
make bugsink-up-exposed the above plus the nginx intake proxy + 8443 (exposed)

make run and make factory do not start nginx or Bugsink. Order does not matter: the proxy retries its upstream, and HydraFlow only reads the DSN at boot. The one thing that must line up is the port — the proxy's upstream defaults to the dashboard's 5555, and a test pins that against config.dashboard_port so the two cannot drift.

The factory runs unattended, so an unreported exception is invisible rather than quiet. The sensor closes the loop from the software failed to the board knows. Two halves, both ours.

1. Start the backend

# .env — the compose file REFUSES to start without these, by design
BUGSINK_SECRET_KEY=$(openssl rand -base64 50)
BUGSINK_SUPERUSER=you@example.com:a-real-password
BUGSINK_DB_PASSWORD=$(openssl rand -base64 24)

make bugsink-up      # http://localhost:8000/  (loopback only)
make bugsink-logs
make bugsink-down    # data survives; the volume is not removed

2. Outbound — point HydraFlow at it

Create a project in the Bugsink UI, copy its DSN, and put it in .env:

SENTRY_DSN=http://<key>@localhost:8000/<project-id>

The DSN's presence is the switch: absent, the composition root returns the no-op adapter, which is what tests, CI and the air-gapped sandbox get. The client is the Sentry SDK, so this same setting can point at sentry.io instead — Bugsink is the house default because self-hosting keeps error payloads from a repo full of unreleased work on infrastructure you own.

HYDRAFLOW_SENTRY_DISABLED=1 overrides a configured DSN, always toward off.

Verifying the loop for real

scripts/bugsink_e2e_smoke.sh

Not in CI (needs Docker, pulls two images). It stands the stack up from nothing, creates a project, throws a real exception through the real adapter, checks Bugsink grouped it, drives the nginx lane, and proves a client cannot choose its own ?source=.

It exists because the unit, config and scenario layers all passed while the loop had never actually been run once — and the first real run found two defects those layers could not see: compose interpolating every service's variables (so a local-only start demanded TLS certs), and the proxy's upstream pointing at Bugsink's port instead of the dashboard's.

The proxy lives in its own docker-compose.intake-proxy.yml for exactly that first reason. Giving its tokens empty defaults instead would have been worse than it looks: an empty HF_EXCEPTION_PATH_TOKEN renders location = /exception/, a real reachable path. A missing token must stop the stack, not open a lane.

Configuring an app that is not HydraFlow

docs/standards/exception_sensor/RUNBOOK.md carries the SRE runbook: SDK init for Python and Node, the check at each step, how Bugsink's grouping differs from Sentry's, and a symptom table for "we get no error issues". Read it before wiring a managed repo.

It lives beside the standard deliberately. .claude/ is stamped into nothing, so a runbook there would have reached HydraFlow and no one else — a repo onboarded to the format would inherit the exception-sensor rules with no instructions for satisfying them. Everything under a kernel standard's directory ships, and test_the_exception_sensor_runbook_ships_with_the_kernel pins that.

How Bugsink groups (it is not Sentry's algorithm)

Bugsink aggregates events into issues, but keys on exception type + message with variable-ish parts normalised (UUIDs, hashes, numbers, URLs, emails, dates), where Sentry keys primarily on the stack trace. Two consequences, both worth knowing before reading the board:

  • The same type and message raised from two different code paths collapses into one issue. Sentry would have split them.
  • A bug whose stack varies stays one issue. Sentry would have split that too, and for a factory this is usually the better failure.

Custom fingerprints are supported ("{{ default }}" includes the automatic grouping and refines it) if under-splitting ever bites.

This is also why the intake deduplicates on the Bugsink issue id rather than the rendered message: the backend has already decided what one group is, and re-keying on text here would second-guess it — and split a group whose message Bugsink deliberately normalised.

3. Inbound — point it back at HydraFlow

Bugsink has no GitHub integration. Its only outbound path is a custom webhook, and that webhook's config is a bare URL: no signing secret, no auth header. So it cannot present the operator credential the intake boundary requires, and the stack ships a proxy that adds one.

# ADR-0140's operator token — the SAME credential the policy write plane uses.
HYDRAFLOW_OPERATOR_TOKEN=$(python -c "import secrets; print('hfop_' + secrets.token_urlsafe(32))")
# One path token per lane, so they rotate independently.
HF_EXCEPTION_PATH_TOKEN=$(openssl rand -hex 32)
HF_REPORT_PATH_TOKEN=$(openssl rand -hex 32)

# then in Bugsink: Alerts -> Custom webhook ->
#   https://<host>:8443/exception/<HF_EXCEPTION_PATH_TOKEN>

Two isolated lanes, one method

The application keeps one intake handler. The separation is at the proxy:

Lane External path Becomes Default rate
Exceptions /exception/<token> ?source=bugsink 30r/m, burst 20
People / UI /report/<token> ?source=ui 10r/m, burst 5

/api/issues/intake is not reachable directly, and that is the point:

  • source is pinned per lane. Passing the intake through untouched would let anyone holding either token file an issue labelled as a system exception — and triage treats those differently, auto-closing the ones that fail. Provenance you can set yourself is not provenance.
  • The tokens rotate independently. Re-keying Bugsink does not disturb the UI, and a leaked report token buys nothing on the exception lane.
  • Separate rate limits. An error storm cannot exhaust the budget a person needs, and neither lane can flood the board on its own.

Remote deploys — expose the intake, never the dashboard

make bugsink-up-exposed adds an nginx front door (docker/hydraflow-proxy/). It is default-deny and opens exactly two lanes, both landing on the one intake method: /exception/<token> and /report/<token> (see below).

Everything else returns 404, and that is not conservatism — it is required. HydraFlow's dashboard has no in-process authentication: ADR-0138 §D5 records that its operator boundary is the loopback bind, and of ~169 dashboard routes, 8 mention the operator token. Publishing the dashboard would expose the control plane, not the UI.

So on a remote deploy:

HYDRAFLOW_DASHBOARD_HOST=127.0.0.1     # keep it. Also what keeps ADR-0140's gate open.
HF_TLS_CERT_DIR=/etc/letsencrypt/live/<host>   # holds fullchain.pem + privkey.pem
HF_SERVER_NAME=hydraflow.example.com
HF_INTAKE_PROXY_PORT=8443              # the ONLY published port

If you ever find yourself setting HYDRAFLOW_DASHBOARD_HOST=0.0.0.0 to make something reachable, that is the signal to put it behind this proxy instead — the intake gate will 404 rather than let the bind go public, deliberately.

Keep the hfop_ prefix: src/secret_scrub.py (ADR-0085) redacts that grammar from the audit, transcript and event streams, and a token without it is only redactable when printed next to its variable name.

The intake boundary is generic — the UI files through the same door with ?source=ui — and it does not exist at all unless the dashboard is bound to loopback. That is ADR-0140's rule, not a new one: a credential checked on an interface the world can reach is a credential the world can brute-force.

What arrives on the board

One issue per error group, labelled hydraflow-find (so triage picks it up) and bugsink (so a human reading the board can tell an observed failure from an authored finding). Bugsink re-fires the same group on regression and unmute; the receiver deduplicates on the group id, so those land as no-ops rather than duplicates. A sensor issue that fails triage is auto-closed as a transient rather than parked for clarification no author will supply.

Onboarding a foreign managed repo

The first foreign managed repo is T-rav/poop-scoop-hero (PSH, a Phaser.js game). Onboarding flow:

  1. Clone the foreign repo locally (git clone git@github.com:T-rav/poop-scoop-hero.git ~/projects/poop-scoop-hero).
  2. Register with HydraFlow's runtime registry:
    curl -X POST http://localhost:8080/api/repos/add \
      -H 'Content-Type: application/json' \
      -d '{"path":"/Users/travisf/Documents/projects/poop-scoop-hero"}'
    
    This validates the path, detects the slug from the origin remote, calls register_repo_cb (→ RepoRuntimeRegistry.register() + RepoRegistryStore.upsert()), and creates HydraFlow lifecycle labels on the repo via ensure_labels.
  3. Add the slug to HYDRAFLOW_MANAGED_REPOS:
    export HYDRAFLOW_MANAGED_REPOS='[{"slug":"T-rav/poop-scoop-hero","main_branch":"main"}]'
    
    This makes PrinciplesAuditLoop audit the repo on its weekly tick. The audit produces a pending → ready (or blocked) onboarding status.
  4. (Optional) Start a RepoRuntime for the repo via POST /api/runtimes/{slug}/start. The runtime runs the orchestrator-style five-loop set in-process. Recommend waiting until the principles audit gives the repo a ready status before flipping this on.

Architectural note (April 2026): ADR-0009 (Accepted) specifies a subprocess-per-repo model with a TCP supervisor (hf_cli/supervisor_service.py). That code lives in a worktree snapshot and was never merged onto main. The in-process RepoRuntime is the working path; isolation (state, event bus, worktree paths) is enforced via per-slug data paths but the Python interpreter is shared. Acceptable at 2 repos. Re-landing the supervisor is a separate ADR-0009 closeout.

Advisor failures degrade per-role; never crash the review

Each advisor role (pre-flight planner, mid-flight consultant, post-verify gate) has a failure-soft contract: pre-flight failure → returns None and the executor proceeds plan-less; mid-flight failure → executor proceeds with its own judgment; post-verify failure → APPROVE by default (FAIL_AS_VETO=true flips the default to VETO for high-stakes surfaces). All advisor runners call reraise_on_credit_or_bug(exc) per §2.2 above, so CreditExhaustedError and likely-bug exceptions still propagate; only "advisor unavailable" failures degrade. The _extract_json_block helper (T16.5 fix) parses production transcripts, which aren't bare JSON — without it the advisor would silently degrade-to-APPROVE on every PR, exactly the failure mode the pattern exists to prevent.

Why: A self-repairing review can't itself become a single point of failure; failure-soft on advisor crashes preserves pre-advisor pipeline behavior. See ADR-0059.

Advisor self-modification guard forces VETO authority on diffs to advisor's own files

When a diff modifies src/review_advisor.py or src/review_phase.py, resolve_post_verify_authority (T29) forces post-verify authority to VETO regardless of surface configuration — including wiki_ingest's normally-advisory mode. The single source of truth for these paths is review_advisor.SELF_MODIFYING_PATHS (frozenset). The wiki_ingest path's _build_wiki_ingest_diff_descriptor synthesizes pseudo diff --git headers when ingest content references advisor source files, so the same substring detector fires for content-driven self-modification, not just file-level diffs.

Why: Prevents the advisor from silently approving changes to its own implementation. Operators encountering "advisor blocked my wiki ingest about advisor source files" should know it's intentional fail-closed behavior. See ADR-0059.