Engineering · Resilience
Design for interruption, ambiguity and recovery.
Reliable software should remain understandable when responses are lost, work is interrupted, state becomes stale, dependencies fail, or an operation completes only partially.
Core practices
- Bounded retry: retry only when the operation and evidence make retry safe.
- Idempotency and duplicate-effect protection: distinguish “no response” from “no effect.”
- Recovery-first state: preserve enough durable state to continue or explain an interrupted operation.
- Stale-state rejection: newer authoritative state must not be silently replaced by older evidence.
- Fail closed: missing or conflicting authority narrows capability rather than silently expanding it.
- Observability: surface what happened, what did not run and what remains unresolved.
Resilience is not hiding errors
A resilient product can still fail. The difference is that the failure is bounded, visible, recoverable where possible, and does not silently manufacture success.
For the open-source project's effect/retry model, see the agent timeout article and CrashLab.