Contract 12
Testing
A robustness library is only as good as the evidence that its guarantees hold, and the code built on it is only as good as the tests its users can write. Both depend on the same thing: making time, randomness and failure controllable.
What StoicTesting provides
| Tool | For |
|---|---|
TestClock |
Anything that races work against time. Moves only when the test moves it; waitForSleepers(count:) is the exact way to know woken work has settled. |
ImmediateClock |
Recording the delays a policy chose. Every sleep returns at once — so never for timeouts. |
withTestAmbient(clock:seed:events:) |
One call that installs a clock, a seed, recorded events and fresh retry budgets. |
@Test(.stoic(seed:)) |
The same isolation as a Swift Testing trait, with StoicTest.events. |
RecordingInstrumentation |
Asserting on decisions, not only results: events.names, events(named:). |
FaultInjector |
Scripted failures, latency, hangs, and hangs that ignore cancellation. |
TestGate |
Sequencing steps, and modelling work that cannot be interrupted. |
StateGraph.explore |
Exhaustive exploration of a state machine's transitions. |
simulate(seed:), simulate(seeds:) |
Deterministic simulation: every interleaving decided by a seed, virtual time, deadlocks and livelocks reported, a failing seed replayed exactly. |
Failure-mode coverage
Every contract's "What happens if…" table is a list of promises, and a
promise without a test is a defect. A test names the row it proves —
@Test(.failureMode("design/09", "a child stops beating")) — and
FailureModeCoverageTests parses every table in design/ on every run:
- a row no test claims fails the suite, unless it is listed in
design/failure-modes.baseline; - a baseline row that a test now claims fails the suite until it is removed, so the baseline only shrinks;
- a claim naming a row that no longer exists — rewritten, misspelt — fails, so a contract cannot drift away from its tests.
STOIC_WRITE_FAILURE_BASELINE=1 swift test --filter FailureModeCoverage
rewrites the baseline from the current state.
Deterministic simulation
FoundationDB tested its database for years in a simulator before trusting real machines with it, and TigerBeetle and Antithesis made the method famous: run the whole system on one thread, let a seed choose every ordering, make time virtual, and a failing seed becomes a recording of the bug. Concurrency bugs live in rare orderings; a stress run reaches them once in thousands of tries, a simulation explores tens of thousands of small orderings a second and replays any of them on demand.
try await simulate(seeds: 0..<500) { sim in
let engine = SyncEngine(store: FakeStore())
try await engine.syncAll()
try #require(engine.pending.isEmpty) // #require: a failing #expect does not stop the seed
}
// simulation with seed 137 threw after 412 steps, 31.5 seconds of virtual
// time: … Reproduces: simulate(seed: 137)How it works. A TaskExecutor queues every job instead of running it;
a driver on one dedicated thread takes the next job from among all runnable
ones by the seed, and runs it. When nothing is runnable, it moves the
TestClock to the next deadline and wakes its sleepers — so time is
virtual and an hour of backoff takes milliseconds. The body and its
structured children run there by executor preference, actors without their
own executor follow the task calling them, and every task Stoic starts
follows Ambient.taskExecutor (unstructured tasks do not inherit a
preference, so Stoic passes it on).
What it reports.
threw: the body failed.simulate(seeds:)runs the failing seed again and says whether it reproduced at the same step.stalled: nothing runnable, nothing asleep, and nothing arrived from outside — a deadlock, or a wait on something the simulation does not drive. Lists the Stoic tasks still running, by label. When no job has ever been enqueued from another thread and noabandoningAfterwork is outstanding, nothing outside is likely to take part, and the stall is reported afterlimits.quietStall(50 ms) rather thanlimits.stall(two seconds of real time, which a hop to a busy main actor may need).virtualTimeLimit: still going afterlimits.virtualTime— a livelock, such as a retry without a bound or a timer that re-arms itself.stepLimit,cancelled.
A run that succeeds returns the value with its step count, virtual time,
the schedule it followed, a fingerprint of every scheduling decision (equal
fingerprints, equal paths) and the Stoic tasks still running when the body
returned (leftover) — almost always a leak. After a failure the body is
cancelled and its cleanup still runs simulated, bounded.
Teardown. When a run ends, success or failure, the simulation cancels
the Stoic tasks of its ledger tag and everything asleep on its clock, and
drives the executor until it is quiet (at most 20,000 steps, a few rounds;
not counted in steps or the fingerprint). Without it a task parked on the
clock kept the clock, and itself, alive for the rest of the process, and a
cancelled one never ran its cleanup. leftover still means what was
running when the body returned; stuck is what was still running after
the unwinding — waiting on a continuation nothing will resume, or ignoring
cancellation. Tasks that are in neither the ledger nor on the clock (a
Task {} of the app's, parked on a continuation) cannot be reached.
Schedules. Which runnable job goes next is a strategy, chosen by the
seed (simulate(schedule:) forces one, and Simulated.schedule and
SimulationFailure.schedule say which ran):
- Uniform: every runnable job equally likely. Tasks stay in lock-step, so the bugs that need two tasks to interleave once — a lost update, a check and act across an await, an ABA, a missed wakeup — are found in a few seeds. It cannot find a bug that needs one task to run many steps while another waits at one point: that takes 2⁻ⁿ.
- PCT (probabilistic concurrency testing, Burckhardt et al.), depth two
to four: every task gets a random priority, the highest-priority runnable
job runs, and at
depth - 1random picks the task that just ran is demoted below everyone. A bug that needsdordering constraints is found with probability at least 1/(n·kᵈ⁻¹) forntasks ofksteps, however long the run-ahead is. The runtime has no API for which task a job belongs to; the task id is read fromUnownedJob.description, in one place that falls back to uniform choice if the format ever changes. A job that is not a task's own — an actor draining its mailbox, whatTask.yieldqueues to resume its task — takes the priority of the task whose job queued it. The run's length is not known in advance, so the change points fall within a horizon the seed picks, sixteen steps to eight thousand. A task that re-enqueues itself 256 picks in a row (a spin-wait onTask.yield) is demoted below everything, or the task it waits for would never run; demoting at the first yield would turn PCT into round-robin and lose the run-ahead bugs. - The default portfolio gives half the seeds to uniform and the rest, in turn, to PCT at depths two, three and four, so a contiguous seed range covers all of them.
Measured over 1,000 seeds, with unrelated suspension points around the critical section (n tasks, about 400 steps at the worst row):
| bug | uniform | portfolio | pct, depth 2 |
|---|---|---|---|
| lost update, no padding / 50 steps | 746 / 117 | 488 / 76 | 23 / 5 |
| actor reentrancy, no padding / 50 | 937 / 230 | 621 / 148 | 51 / 6 |
| ABA across two awaits, 50 | 22 | 10 | 2 |
| missed wakeup, 50 | 64 | 42 | 2 |
| 3-lock deadlock, 50 | 24 | 15 | 4 |
| one task 10 yields ahead of another's one | 0 | 143 | 405 |
| … 40 yields ahead | 0 | 97 | 297 |
The portfolio keeps about two thirds of uniform's hits on the classics and finds what uniform never does. PCT alone is worse on the classics, which is expected: it runs one task ahead unless a change point lands exactly in the window.
Early timers. Virtual time otherwise passes only when nothing is
runnable, so work always beats a timeout: "the timer fires in the middle of
the computation" could never be explored. On half the seeds (chosen
independently of the strategy) the next deadline may also fire while jobs
are still runnable, with probability 1 in 8, 32 or 128 at each step and at
most 32 times per run (so progress is never starved), and never while
abandoningAfter work is outstanding. Earliest deadline first, so order
among timers is still causal. A 1 ms timeout against ten steps of work: 0
of 1,000 seeds without early timers, 252 with. schedule: .uniform (or .pct(depth:)) turns them off; .earlyTimers(true) turns
them on.
Fault points (buggify). Scheduling explores orderings; it cannot make a
finalizer slow or a retry attempt late. So Stoic's primitives carry named
fault points where timing decides the path — semaphore.arriving,
singleflight.arriving, channel.send, slot.deliver, scope.finalizer,
statemachine.work, service.start, supervisor.child, retry.attempt.
Under simulation (faults: .standard) each site is switched on for a
quarter of runs, chosen by the seed, and fires a quarter of the times it is
reached: a few yields, or up to a second of virtual time. faultsFired
says which fired. Outside a simulation a fault point is one relaxed atomic
load. The test that proves it: a scope whose finalizer has a 100 ms
timeout passes 200 seeds without faults and fails within 200 with them,
reproducibly — the finalizer was slow to start and was abandoned, so the
resource was still open after the scope returned.
Reviewed. An adversarial review found, and the tests now pin: the body
did not inherit the caller's task-locals (it is now started in the
caller's context); a perpetually runnable task left behind turned a
finished body into a step-limit failure (the body's end now ends the run);
virtual time raced ahead of work handed to Dispatch (time now stands still
while abandoningAfter work is outstanding, until it proves wedged); and
primitives cancelled sets of tasks in dictionary order, which changes with
each process's hash seed, so a seed did not replay across processes (they
now go in key order, and the fingerprint includes each job's arrival
number, so a different path cannot hide behind the same fingerprint).
Failures inside the body must throw: use #require, or throw — a failing
#expect is recorded but does not fail the seed.
Cost. About three million steps a second, 30 µs for a tiny run (release
build). PCT reads each job's description, so a PCT run is several times
slower per step. A timer heap in TestClock makes 30,000 sleepers with
distinct deadlines fire in 0.11 s (19.7 s when each fire scanned them all).
Limits, stated rather than hidden. The main actor and actors with their
own executors run where they always do; so do Task {} started by app code
(start it from a Stoic primitive, or with Task(executorPreference: Ambient.current.taskExecutor)), Dispatch, real I/O and real clocks. Work
arriving from there is waited for, but its timing is not seeded, so such a
run may not replay exactly; the fingerprint says when it did not. A job
that blocks its thread blocks the simulation.
Verified. The same seed gives the same trace, step count, virtual time and fingerprint; twenty seeds give many distinct interleavings of a scenario with a task group, an actor, a semaphore and a jittered retry; a read–yield– write lost update is found within 200 seeds and reproduces; a deadlock is reported with the stuck state work by name; a livelock hits the time limit.
How Stoic tests itself
- Every failure-mode row has a test, on virtual time where time is involved.
- Properties for arithmetic and state. Schedules over 300 generated configurations; seeded jitter reproducibility; budget accounting; schema round trips and sparse fixed points over hundreds of generated values; an exhaustive 255-combination accumulation test; a differential check of the JSON Schema export against an independent validator.
- Stress. Every concurrency suite must pass
swift test --filter <Suite> --maximum-repetitions 300 --repeat-until fail. This found three real races during development (a cancellation cause recorded too late, a semaphore waiter cancelled between deadline and enqueue, a ledger shared across repeated runs) — each fixed at its cause, never by a sleep. - Process-wide state is isolated per test (budgets in the ambient context; unique labels for process-wide ledgers), so tests can repeat and run in parallel.
- No real sleeps as synchronisation. Tests wait on clocks, gates and results.
Planned
- A task ledger and
assertQuiescent(). Every Stoic-owned task registered with its label and owner; a test's teardown asserts nothing is left running, no sleeper is pending, no scope is open and no work is abandoned — leaks become test failures that name their owner. - Adaptive horizons. PCT guesses the run's length from the seed; a
simulate(seeds:)that learns it from the first passing runs would find more with the same seeds, at the price of a replay recipe that names the horizon too. - More fault points. Store writes torn or failing under simulation (the simulated file system already can), and a policy that targets the sites a failing seed hit, to shrink it.
- Smarter scheduling. Uniform choice finds shallow bugs fast; PCT-style priority schedules find deep ones with known probability. And shrinking: replay a failing seed with fewer preemptions until the minimal ordering remains.
- Generated test skeletons from the failure-mode tables, for a new contract.