Contract 09
Supervision
Long-lived work fails. A sync engine loses its connection, an event tap
gets disabled by the system, a model server crashes on a bad input. The
hand-rolled answers — a watchdog timer that pings and aborts, a while true { do { … } catch { sleep } } loop — restart too eagerly or not at all, fire
during a slow launch or right after the Mac wakes, and turn a persistent
fault into a crash loop. Supervision is the discipline for this, and Erlang
has had the specification right for thirty years.
Origin
- Erlang/OTP supervisors. Restart strategies (
one_for_one,one_for_all,rest_for_one), restart types (permanent,transient,temporary), an intensity window that escalates when restarts are not helping, ordered start and reverse-ordered shutdown with per-child timeouts. - Akka. Backoff with randomisation between restarts.
- ServiceLifecycle. A ladder from graceful shutdown to cancellation to giving up, and no restart option at all — the gap this fills.
API
public final class Supervisor: Sendable {
public init(label: String, strategy: Strategy = .oneForOne, intensity: Intensity = .standard,
backoff: Schedule = .exponential(base: .milliseconds(500)).capped(at: .seconds(30)),
resetAfter: Duration = .seconds(30), @ChildrenBuilder children: () -> [Child])
public func run() async throws // SupervisorEscalation, or CancellationError
public var events: AsyncStream<SupervisorEvent>
}
public struct Child { init(_ label: String, restart: RestartPolicy = .permanent, heartbeat: HeartbeatPolicy? = nil,
shutdownTimeout: Duration = .seconds(5), run: @escaping @Sendable (Heartbeat) async throws -> Void) }
public enum RestartPolicy { case permanent, transient, temporary }
public struct HeartbeatPolicy { init(interval: Duration, misses: Int = 3, grace: Duration = .seconds(30)) }
public final class Heartbeat { public func beat() }Semantics
One loop decides. Children, heartbeat monitors and restart timers send signals to a single loop that owns all state. Every child run has a generation; a signal from a generation that has been stopped or replaced is stale and ignored, so a child stopped by the supervisor can never be mistaken for one that crashed.
Restarts. permanent children restart whenever they stop; transient
only when they throw; temporary never. When every child has finished for
good, run() returns.
Strategies. oneForOne restarts the stopped child. oneForAll stops the
others newest first and restarts all in order. restForOne does the same
for the stopped child and those started after it.
Intensity. More than maxRestarts within window (five a minute by
default) means restarting is not helping: every child is stopped and run()
throws SupervisorEscalation with the child and its last error, for the
next level up — a parent supervisor, a service graph, or the process. A
crash (a trap) cannot be caught in-process; that level is the operating
system restarting the process, and is configured there (launchd's
KeepAlive, the app being relaunched).
Overlapping restarts. A newer restart decision supersedes any restart already pending for the children it covers, and a child never runs two instances: so a second failure during a group's backoff cannot start members out of order or twice. A child that finished for good (a transient that returned, a temporary) stays finished when its group restarts.
Children own their time. A child runs without the deadline of whoever
called run(): it lives as long as the supervisor keeps it. Restart delays
are counted from the decision and heartbeat checks fall on fixed instants
from the child's start, so neither a timer's own scheduling delay nor drift
across intervals stretches them.
Backoff. Each child waits on the supervisor's schedule before
restarting (jittered exponential by default, capped at 30 s); a child that
ran for resetAfter without stopping starts its backoff over. A schedule
that ends means restarting should stop: the supervisor escalates.
Heartbeats. A child with a HeartbeatPolicy calls beat() while it
makes progress. Silence for misses consecutive intervals (three by
default — one slow moment is not a hang) means hung: the child is stopped
and restarted like a failure. Heartbeats use the liveness clock, which
stops while the machine sleeps, and are not checked during grace after a
start, for slow launches. Both rules exist because watchdogs that ignored
them have turned slow launches and wake-from-sleep into crash loops.
Stopping. Newest first. Each child is cancelled with cause
supervisorStopped and given its shutdownTimeout; one that does not stop
is abandoned and reported (supervisor.child_abandoned) so shutdown always
finishes.
Failure modes
| What happens if… | Behaviour |
|---|---|
| a permanent child throws | Restarted after its backoff |
| a transient child returns | Done; not restarted |
| a temporary child throws | Done; reported |
| restarts exceed the intensity | All children stopped; SupervisorEscalation thrown |
| a child stops beating | After misses intervals (past grace): stopped and restarted |
| the Mac sleeps | The liveness clock stops; no false hang |
| the supervisor's task is cancelled | Children stopped newest first; CancellationError |
| a child ignores cancellation on shutdown | Abandoned after shutdownTimeout; reported |
| a child ran long and then failed | Its backoff starts over |
Events
supervisor.child_started, supervisor.child_exited (failed or not, with
the error), supervisor.child_restarting (delay), supervisor.child_hung,
supervisor.child_abandoned, supervisor.escalated, supervisor.stopped;
also published on events for a UI or a parent.
Testing
Each row above, on the virtual clock: restart after backoff, transient and temporary completion, escalation, both multi-child strategies, a silent child detected and a beating one left alone, cancellation order and causes, and abandonment. 300 repetitions are clean.
Planned
Supervisoras a ServiceLifecycleService(trait), and as a child of a parent supervisor without ceremony.- "Significant" children whose completion stops the supervisor (OTP's
auto_shutdown). - A crash-loop marker for the process level: unclean launches counted across restarts, with a safe mode after too many (design/15, planned).