Contract 02
Schedules and retry
Retries are policy, not loops. A Schedule is a value that decides, attempt
by attempt, whether to go again and how long to wait. retry and recur
run an operation under one. Budgets make sure that retrying can never turn an
outage into a bigger one.
Origin
- Effect.
Scheduleas a composable value: exponential, spaced, recurs, union and intersection, jitter. - AWS Builders' Library. Five layers each retrying three times multiply
load on a failing database 243-fold. Retry at one layer; cap retries with a
local token bucket; use capped exponential backoff with jitter; retry only
idempotent operations; never retry caller errors; honour the server's
Retry-After. - Finagle, gRPC, Envoy. Retry budgets as a fraction of recent requests, not a per-call count.
- In apps. Retry loops with a fixed count and a fixed delay, no jitter, and no awareness of cancellation or of how much time the caller has left. Their numbers are written in the code, the comments and the docs, and drift apart.
API
public struct Schedule: Sendable, Hashable { // data: an inspectable tree, not closures
public static func fixed(_:) / linear(_:) / exponential(base:factor:) / fibonacci(_:)
public static func decorrelated(base:cap:)
public static func delays(_ delays: [Duration]) -> BoundedSchedule
public func capped(at:) / jittered(_:) / upTo(elapsed:) / union(_:) / intersection(_:) / andThen(_:)
public func attempts(_ count: Int) -> BoundedSchedule // the only way to bound a schedule
public func maximalDelays(limit:) -> [Duration]
}
public struct BoundedSchedule: Sendable, Hashable { // what retry accepts
public var maximumAttempts: Int
public var maximumTotalDelay: Duration
public func worstCase(attemptTimeout: Duration) -> Duration
}
public enum RetryDecision { case retry, retryAfter(Duration), stop }
public protocol Retryable: Error { var retryDecision: RetryDecision { get } }
public struct RetryClassifier<Failure: Error> { static var standard, transientOnly, never; init(_:) }
public enum RetryNesting { case singleAttempt, allow }
public struct RetryBudget { static let perCallSite, unlimited; static func named(_:); init(ratio:reservePerSecond:reserveCapacity:) }
// Inline: runs on the caller's actor, may capture anything.
public nonisolated(nonsending) func retry<R, Failure: Error>(
_ schedule: BoundedSchedule, budget: RetryBudget = .perCallSite,
classify: RetryClassifier<Failure> = .standard, nesting: RetryNesting = .singleAttempt,
label: String? = nil, operation: nonisolated(nonsending) (Attempt) async throws(Failure) -> R
) async throws(Failure) -> R
// Each attempt bounded by its own timeout; runs as a child task, inheriting the caller's actor.
public nonisolated(nonsending) func retry<R: Sendable, Failure: Error>(
_ schedule: BoundedSchedule, attemptTimeout: Duration, …,
@_inheritActorContext operation: @escaping @Sendable @isolated(any) (Attempt) async throws(Failure) -> R
) async throws(Failure) -> RA call site with its whole budget stated once, and the outer watchdog derived from it rather than written down separately:
let start = Schedule.fixed(.milliseconds(650)).attempts(2)
try await withTimeout(start.worstCase(attemptTimeout: .milliseconds(2500))) {
try await retry(start, attemptTimeout: .milliseconds(2500)) { _ in try await microphone.start() }
}Semantics
Bounded by type. retry takes a BoundedSchedule, and the only ways to
make one are attempts(_:) and delays(_:). An unbounded retry does not
compile; there is no runtime cap that behaves differently in debug and
release. upTo(elapsed:) is an extra stop condition, not a bound: a schedule
with no delay and only an elapsed limit could spin.
Schedules are data. A schedule is a tree of nodes interpreted by a
stepper, so it prints, compares, reports its worst case, and steps the same
way in a test as in production. Full and equal jitter only ever shorten a
delay, so a cap holds whichever order capped and jittered are written
in. capped(at:) directly after a base replaces the base's default cap
(exponential and Fibonacci default to five minutes).
Order of checks after a failed attempt. (1) The task running the retry
was cancelled → rethrow; a cancelled attempt (its own timeout fired) is not
a cancelled retry. (2) Nested inside another retry's attempt with
.singleAttempt → rethrow, retry.nested. (3) The classifier says stop (a
Defect always does) → rethrow. (4) The schedule is exhausted → rethrow.
(5) The next attempt could not finish before Deadline.current, assuming it
takes as long as the last → rethrow, retry.deadline_skip. (6) The budget has
no token → rethrow, retry.budget_exhausted. Otherwise sleep (cancellable)
and try again. The error is always the last attempt's, unwrapped; cancelled
during the backoff, retry rethrows that last real failure.
Default classification. .standard lets Retryable errors decide,
stops on defects, DecodingError and EncodingError (the same bytes fail
the same way), and on rejections retrying cannot help (an exhausted budget,
a shutdown, a wedged label), honours a rejection's retryAfter, treats a
timed-out attempt as transient, and retries anything else. Review argued for
"transient only" as the default; the counterweight is that every other
safety net is on by default — the schedule is bounded by type, the budget is
per call site, nested retries collapse — so an unknown error costs a few
bounded attempts, not an outage. .transientOnly is one word away.
Nesting. A task-local marks a running attempt. A retry started inside
one makes a single attempt by default (retry.nested), enforcing AWS's
"retry at one layer" at runtime instead of documenting it. .allow opts out
for layers that retry different dependencies.
Budgets. On by default, one per call site. Each first attempt deposits
ratio (0.2) earned tokens; a reserve refills at reservePerSecond (1) up to
reserveCapacity (10) so a quiet app can still retry; each retry spends one.
Under a full outage, retries add at most 20% load plus one per second per
call site. Named budgets (.named("api.example.com")) share one bucket
across call sites hitting one dependency. Earned tokens are capped so a long
success streak cannot bank a burst for the next outage.
Arithmetic. Delays saturate instead of trapping (Duration
multiplication traps on overflow), so an exponential at attempt 300 is just
its cap.
Jitter and seeds. Full jitter is opt-in per schedule (.jittered()).
Randomness comes from the ambient source, one stream per call site, so a seed
reproduces a run's delays; in production the source is the system generator
and concurrent callers draw independently.
Retry-After. .retryAfter(d) replaces the schedule's delay for that
retry but not its bound, its budget or the deadline, and is capped at five
minutes.
Recur. recur (next) runs an operation on a schedule after success
— polling, periodic refresh — and may be unbounded, but only inside a scope
or supervisor that can stop it.
Failure modes
| What happens if… | Behaviour |
|---|---|
| the operation succeeds on attempt n | Result returned; retry.succeeded with n (only when n > 1) |
| every attempt fails | Last error rethrown; retry.exhausted with attempts and elapsed |
| the task is cancelled during the backoff sleep | The last attempt's error is rethrown (typed); no further attempt |
the error is a Defect |
Rethrown immediately; retry.defect at error level |
| the budget is empty | Last error rethrown; retry.budget_exhausted |
| the next attempt cannot finish before the deadline | Last error rethrown without sleeping; retry.deadline_skip |
| the server sends Retry-After longer than the cap | The cap wins; retry.after_capped |
| the schedule is unbounded | Does not compile: retry takes only a BoundedSchedule |
| a retry runs inside another retry's attempt | One attempt only, retry.nested, unless nesting: .allow |
| an attempt's own timeout fires | Counted as transient and retried; timed_out=true on the event |
the error is a DecodingError |
Not retried (.standard): the same bytes fail the same way |
Events
retry.attempt_failed (attempt, error type, decision, next delay),
retry.succeeded, retry.exhausted, retry.budget_exhausted,
retry.deadline_skip, retry.defect, retry.permanent, retry.nested,
retry.after_capped.
Testing
- Property tests over random schedules: delays are within
[0, cap], never negative, attempt bounds are exact,union/intersectionobey their laws (commutative,upTomonotone), and saturating math never traps across the whole attempt range. - A seed reproduces the identical delay sequence under randomly interleaved concurrent retries.
- Budget property: over any request sequence, retries ≤ ratio × first attempts + floor × window.
- Virtual-clock tests for every failure-mode row, with fault-injected
operations (
fail(times:),fail(with: Retryable(.stop)),hang).
Non-goals
Retrying non-idempotent operations safely (the caller's responsibility, with
idempotency keys); retry inside AsyncSequence operators (async-algorithms'
domain).