Failure injection
This page covers Phase 13 (P13-05 to P13-12) and AGENTS.md section 42. Each fault must have an observable recovery path. No fault may silently corrupt the Canonical State.
Two kinds of proof exist:
- Critical tests run the real decision code with the fault applied. They run in
bun run test:criticaland in CI. - Live injection (
bun run inject) applies faults through the public dev API, as the Clerk smoke user, and reports pass or fail per fault.
The live run does not inject a fault that needs an internal hook. This page lists each needed hook. No one has added these hooks yet. Each hook is a test-only path into production code and needs its own review.
Run live injection
Section titled “Run live injection”flowchart LR
start[bun run inject] --> sel{Flags}
sel -->|default| d[API faults:<br/>retry, double submit, budget]
sel -->|"--sandbox"| c[Candidate faults:<br/>one Sandbox verification]
sel -->|"--only=release-smoke-failure"| r[Release smoke fault]
d --> rep[pass or FAIL per fault]
c --> rep
r --> rep
rep --> keep{"--keep set?"}
keep -->|no| arch[Archive created Projects]
keep -->|yes| stay[Keep Projects]
bun run inject # default faults: no Sandbox, no model calls, ~15 D1 rowsbun run inject --sandbox # adds the Candidate faults: one Sandbox verificationbun run inject --only=stale-candidate,promotion-double-submit --sandboxbun run inject --keep # do not archive the Projects it created- Credentials: The same Keychain items as
bun run smoke:dev(rhumbatron-clerk-secret-key-dev,rhumbatron-clerk-smoke-user-id-dev). No new secret. - Target: Dev only (
https://api-dev.rhumbatron.com). Dev must be running (not paused). - Model calls: None. The script creates every Change with
startPlanning: false. - Sandbox: None by default.
--sandboxruns one Acme Shop baseline Candidate verification. It uses a few minutes of the 5 h monthly dev Sandbox cap (ADR 0017). The stale Candidate stops at workflow step 1 and never opens the Sandbox. - Cleanup: The script archives every Project that it created with
DELETE /v1/projects/:id. The archive deletes the Project’s Artifacts repos. Evidence and D1 rows stay. - The script reports a fault that could not run as
FAIL, never as a skipped pass (section 34.4.1).
| Fault ID | Default | What it does | Expected recovery |
|---|---|---|---|
change-retry |
yes | Creates a Change twice with one clientRequestId. |
201, then 200 with the same Change ID. |
change-double-submit |
yes | Sends the two creates concurrently. | Both succeed with one Change ID; the Project stores one Change. |
project-double-submit |
yes | Creates two Projects with one slug concurrently. | Exactly one 201 and one 409. |
budget-limit |
yes | Creates Changes with budget $1000 and $0. |
Both 400. |
candidate-without-source |
yes | Asks for a Candidate on a Project with no source. | 409; no Candidate row; no Sandbox. |
candidate-double-submit |
--sandbox |
Three concurrent Candidate creates for one composition. | All 202, one Candidate ID, one Candidate row, one verification. |
early-promotion |
--sandbox |
Promotes while the Candidate still verifies. | 409 blocked; the Canonical Generation does not move. |
approval-cross-change |
--sandbox |
Approves a Candidate through another Change’s path. | 404; no decision recorded. |
approval-double-submit |
--sandbox |
Two concurrent approvals, then a conflicting reject. | Without a gate: both 409 not_required. With a gate: one decision, identical repeat accepted, reject 409 already_decided. |
promotion-double-submit |
--sandbox |
Two concurrent promotions of the verified Candidate. | Both 200; exactly one first promotion; generation advances once. |
stale-candidate |
--sandbox |
Promotes a Candidate built on the old base. | 409 stale; Candidate marked stale; generation does not move. |
release-smoke-failure |
--only |
Probes the demo Project’s latest deployed release with the normal smoke probes plus a route that does not exist (P13-12). | 200; release smoke_failed with release-s<N>-injected Evidence; a repeat is repeated; generation does not move. |
Fixed risk found while writing change-double-submit
Section titled “Fixed risk found while writing change-double-submit”POST /v1/projects/:projectId/changes read by clientRequestId and then inserted. Two concurrent
requests could both miss the read. The second insert then failed with 500 on the UNIQUE (project_id, client_request_id) constraint (migration 0001). workers/api/src/projects.ts now
inserts with ON CONFLICT (project_id, client_request_id) DO NOTHING and reads the row back.
Both callers get the same Change (section 60).
Fault map
Section titled “Fault map”| P13 | Fault (section) | Proven by critical tests | Live (bun run inject) |
Gap or hook needed |
|---|---|---|---|---|
| P13-01 to 04 | Budget stop (24, 44) | packages/model-router/src/budget.critical.test.ts (per-run call cap per tier, cost cap, Task share never negative, budget stop ends the Task without escalation); route.critical.test.ts (budget_exhausted stops) |
budget-limit |
See budget stop. |
| P13-05 | Duplicate delivery (42.1, 42.2, 60) | workers/event-router/src/dedupe.critical.test.ts (Event applied once per consumer group); workers/coordinator/src/view-model.critical.test.ts (duplicate Event ID ignored); workers/event-router/src/archive.critical.test.ts (redelivered batch, same R2 key); workers/event-router/src/projector.critical.test.ts (out-of-order generation); workers/agents/src/wake-consumer.critical.test.ts (redelivered and concurrent wake: one assignment, one run); workers/coordinator/src/promotion-cas.critical.test.ts (promotion retry); workers/agents/src/conflicts.critical.test.ts (one conflict from either detector) |
change-retry, change-double-submit, project-double-submit, candidate-double-submit, approval-double-submit, promotion-double-submit |
Live K2 redelivery needs a replay hook. See duplicate Event. |
| P13-06 | Expired Claim (42.3) | workers/coordinator/src/claims.critical.test.ts (expired claim does not block and is reported overdue, renew after expiry rejected, alarm at the earliest lease); packages/protocol/src/activity.critical.test.ts (a lapsed lease shows as expired) |
none | Claims are internal RPC only. See expired Claim. |
| P13-07 | Sandbox loss (30.1, 42.5) | workers/integration/src/sandbox-loss.critical.test.ts (a check after Sandbox loss restores the checkout from the Candidate fork at the composed commit; an intact checkout is reused; a missing commit fails, never checks the wrong tree; repo tokens are revoked) |
none | See Sandbox loss. |
| P13-08 | K2 publication outage (18, 42.6) | workers/coordinator/src/outbox-recovery.critical.test.ts (Events kept during the outage and published once K2 recovers; attempt limit keeps the row visible as failed); outbox.critical.test.ts (capped backoff) |
none | See K2 outage. |
| P13-09 | Model timeout (42.8) | packages/model-router/src/route.critical.test.ts (timeout retries once at the same tier, a second transport failure stops, never escalates); decision.critical.test.ts (Clef abstains on malformed output); workers/integration/src/root-planner.critical.test.ts (planner repairs once, then gives up) |
none | See model timeout. |
| P13-10 | Poison command (20, 42.7) | workers/agents/src/queue-poison.critical.test.ts (a failing wake is retried, never acked; the rest of the batch is acked; invalid and stale wakes are acked); scripts/inject/dead-letter.critical.test.ts (every Terraform and local Queue consumer has a dead-letter Queue and 1 to 5 retries) |
none | The event router consumes the dead-letter Queue (workers/event-router/src/dead-letter.ts, tested in dead-letter.critical.test.ts): the Task becomes blocked and the Project and Change pages show Retry and Dismiss. A live poison hook is still to build. See poison command. |
| P13-11 | Stale Candidate (36, 42.9) | workers/coordinator/src/promotion-cas.critical.test.ts (stale expected generation, changed tree, race loser marked stale); packages/evidence/src/promotion.critical.test.ts (base behind canonical reports stale) |
stale-candidate, promotion-double-submit, early-promotion; also bun run smoke:sandbox |
None. |
| P13-12 | Deployment failure (21.3, 42.10) | workers/integration/src/release.critical.test.ts (one release per generation; a newer promotion supersedes an in-flight release, which can never deploy afterwards; a 404 feature probe or a probe without a response never marks a release deployed; only the build of the promoted commit counts); promotion-cas.critical.test.ts (generations only advance) |
release-smoke-failure (--only) |
See deployment failure. |
Hooks needed
Section titled “Hooks needed”Budget stop
Section titled “Budget stop”OpenRouter free models cost $0. On that path, checkRunBudget receives spentUsd: 0 in
AgentDO.beforeModelCall. The cost cap cannot fire live without spending Vertex credit, so the
live stop is the call cap. The per-run cap fails the Task. The daily cap (AGENT_DAILY_CALL_CAP,
D1 usage_counters row agent-model-calls) ends the run with budget_exhausted. failTask
turns that into Change status budget_blocked.
Hook: A dev-only internal RPC on the agents Worker that sets today’s agent-model-calls counter
to the cap for the inject run. Expected: The first agent run of a planned Change stops with
budget_exhausted. The Change shows budget_blocked, and the existing state stays readable.
This costs at most one Root Planner call.
Of the section 44 limits, workers/agents/src/limits.ts enforces max_active_agents_per_change
(8), max_active_sandboxes_per_change (4), and the attempt limit per tier (2). packages/conflict
enforces maxCandidates (3). Nothing enforces the per-Change max_sandbox_seconds_per_change and
max_browser_seconds_per_change. The monthly Sandbox cap and the daily Browser Run cap per
environment apply instead.
Duplicate Event
Section titled “Duplicate Event”Hook: a replay command for the event router that re-consumes one K2 offset range for one
subscription (for example a k2.replay integration command with stream, subscription, and offset
range). Expected: claimEvents reports every Event as a duplicate. The live view sequence does
not advance, and no second R2 archive object appears.
Expired Claim
Section titled “Expired Claim”Claims are ResourceShardDO RPCs (acquireClaim, renewClaim, releaseClaim) with no public
route. Hook: a dev-only internal RPC that acquires a Claim for a test agent with ttlMs: 5000 and
never renews it. Expected: After the lease ends, the shard alarm emits claim.expired. The
activity feed shows the Claim as expired, and a second agent can acquire the resource.
Sandbox loss
Section titled “Sandbox loss”Hook: a dev-only internal RPC on the agents Worker that calls destroy() on a Project’s Sandbox
(getSandbox(SANDBOX, "proj-<id>")). Run it between Candidate composition and the first check,
then let CandidateVerificationWorkflow continue. Expected: The checks pass on a checkout restored
from the Candidate fork, and the Candidate verifies. Task Workspaces need the same proof.
openWorkspace reuses an existing fork on a retried run, but no test covers that path yet.
K2 outage
Section titled “K2 outage”Hook: a dev-only runtime flag (for example KV key inject:k2-down-until in rhumbatron-config-dev)
that OutboxPublisher checks before each send and treats as a retryable K2 failure. KV values are
runtime data, not Terraform configuration. Expected: Mutations still commit, and unpublished
outbox rows grow. The dashboard lags. After the flag expires, every Event publishes once.
Model timeout
Section titled “Model timeout”Hook: a dev-only per-Change flag that makes the gateway client abort the next model call after
1 ms. Expected: One same-tier retry, and then the Task stops with repeated_transport_failure.
No escalation to a stronger tier occurs. Live injection spends free-tier model calls (50 per day).
Poison command
Section titled “Poison command”Consumers retry a failure. After 3 retries, the Queue moves it to rhumbatron-dead-letter-dev.
The event router consumes the dead-letter Queue (workers/event-router/src/dead-letter.ts). It
records task.failed and blocks the Task. The Project and Change pages then show Retry and
Dismiss (section 42.7).
Still needed for a live proof: a dev-only internal RPC that sends one schema-valid command whose handler always throws (the public API cannot enqueue arbitrary commands, by design).
Deployment failure
Section titled “Deployment failure”ReleaseWorkflow (ADR 0018) observes the Workers Build of each promoted commit, confirms the deployment, and smoke checks the release URL. Injection:
bun run inject --only=release-smoke-failure # dev; RELEASE_PROJECT_ID=<id> to pick a Project- The fault needs a deployed release. That is the smoke user’s
acme-shop-liveProject (bun run demo:live) with its Workers Builds connection and one promotion that deployed. The fault never runs by default, because it changes the demo Project’s release state. - What happens:
POST /v1/projects/:projectId/releases/inject-failureanswers only whereFIXTURES_ENABLEDis set, or in dev and local. It probes the live release URL with the normal smoke probes plus/__rhumbatron/injected-missing-route. The app returns a real 404, so the smoke check fails through the normal path. The result is releasesmoke_failed, eventrelease.smoke_failedwithinjected: true, and Evidencerelease-s<N>-injectedin R2 and D1. - Expected: The Candidate page and the inbox show the failed release. The Canonical Generation stays, and nothing is deleted.
- Recovery is a new canonical transition (section 42.10). Run a new Change and promote it. Its
release (
S<N+1>) builds, deploys, and passes the smoke check. - A real build failure (a promoted commit whose build fails) takes the
failedpath withdeployment.failed. The live run does not inject it, because it needs a broken commit on canonicalmain.