Contract 06
Validation and schema
Parse, don't validate. Data that crosses a trust boundary — a config file, a network response, a document from disk, a deep link — enters the app through one declaration per type, and that declaration is the single source of truth for decoding, encoding, defaults, every error message, the JSON Schema, test generators, UI metadata and migrations. Nothing else repeats it, so nothing can drift from it.
Origin
- Alexis King, "Parse, don't validate." A value that has been checked should have a type that says so; the check and the type are one thing.
- Effect Schema, Zod, Pydantic. One declaration drives decode, encode,
validation, JSON Schema and arbitrary generators. Effect's
toArbitrarymakes property tests of any schema free. - What goes wrong in apps. Ranges and defaults repeated in the parser,
the clamp, the settings UI and the docs. Invalid input silently clamped
instead of reported.
Codablethat stops at the first error and points at no line. Synthesised enum coding that is unusable on the wire. A changed default that silently flips the value of every user who had chosen the old one. Older app versions that destroy fields a newer version wrote.
The rule that makes it credible
Every constraint is data, never a closure. A closure validator is
invisible to the JSON Schema exporter, the generator and the UI. Constraints
are values (.range, .length, .pattern, .oneOf, .multipleOf,
.unique, .custom(id:)), interpreted by the decoder and the exporter
alike, so the check and the documentation cannot disagree. Closures exist
only behind a labelled escape hatch (.custom(id:description:check:)), which
the exporter renders as an annotation and the generator treats as opaque.
Layers
Layer 0 StoicJSON strict parser: duplicate keys, exact numbers, positions, limits (shipped)
Layer 1 Refined<R> parse-don't-validate leaf types; constraints as data
Layer 2 Schema<Value> declaration + decoder + encoder; accumulating Issues with paths
Layer 3 Tooling JSON Schema export · generators · SchemaSuite self-test · migrations
Layer 4 Sugar @Schema / @Refined macros (separate product) · Codable bridgeThe macro is last and purely sugar: it emits calls to the Layer 2 API. A person without macros writes the same schema by hand, and it is just as checked.
Layer 1: refined types
public protocol Refinement: Sendable {
associatedtype Base: Sendable
static var constraints: [Constraint] { get } // data: the only source of truth
}
public struct Refined<R: Refinement>: Sendable {
public let value: R.Base
public init(_ raw: R.Base) throws(ValidationFailure) // the only public way in
public init(clamping raw: R.Base) where R.Base: Comparable // the explicit lossy path, for UI
}
// Integer generic parameters give zero-boilerplate integer ranges:
public enum IntRange<let lower: Int, let upper: Int>: Refinement { … }
public typealias Port = Refined<IntRange<1, 65_535>>Refined checks by interpreting constraints, so the check that admits a
value and the JSON Schema that documents it are the same data.
Layer 2: the schema value
public struct Schema<Value: Sendable>: Sendable {
public var node: SchemaNode // inspectable description
public func decode(_ value: JSONValue, options: DecodeOptions = .strict) -> Decoded<Value>
public func decode(_ bytes: some Collection<UInt8>, options: DecodeOptions = .strict) throws(ValidationFailure) -> Value
public func encode(_ value: Value, style: EncodeStyle = .full) -> JSONValue
public func jsonSchema() -> JSONValue // Draft 2020-12
}
public struct Decoded<Value> { public var value: Value?; public var issues: [Issue] }
public struct Issue: Sendable, Hashable {
public enum Severity { case error, repaired, warning }
public var severity: Severity
public var path: JSONPath // "$.network.hosts[3]"
public var code: IssueCode // stable, extensible: "out_of_range", "unknown_key", …
public var message: String // derived from code and parameters
public var expected: String?
public var actual: String? // never filled for secret fields
public var suggestion: String? // did-you-mean
public var source: SourceRange? // when decoded from text
}
public struct ValidationFailure: Error { public var issues: [Issue] }Accumulation. Decoding never stops at the first problem. Every field is
decoded independently; every issue is collected with its path and, when the
input was text, its line and column. A value is produced only when no issue
is an error (or, under .lenient, after repairs).
Decode options.
.strict— any error fails. For wire formats..lenient— an invalid field falls back to its default and records arepairedissue; for hand-edited configuration, where the app must still launch and tell the user exactly what it ignored.unknownKeys: .reject | .warn | .ignoreper object, with did-you-mean.
Did-you-mean. Candidates are declared keys plus aliases. Case and
snake/camel differences are an exact alias match (scroll_speed →
scrollSpeed, reported as a deprecated spelling, not an unknown key).
Otherwise Damerau–Levenshtein with threshold min(2, length/3), at most
three suggestions, deterministic order. The same machinery serves enum
values and union tags. Type confusion gets its own hint: "10" where a
number is expected says "did you mean 10?".
Objects, two ways.
(a) Config-like types with an init(): the default instance is the single
source of every default, and of sparse encoding.
extension Settings {
static let schema = Schema.object(Settings.self, default: Settings()) {
Field("scrollSpeed", \.scrollSpeed, .double(.range(200...4000)), doc: "Pixels per second.", unit: "px/s")
Field("theme", \.theme, .enumeration(Theme.self))
Field("hosts", \.hosts, .array(.string(.length(1...253)), count: 0...8), aliases: ["scroll_hosts"])
Field("apiToken", \.apiToken, .string(), secret: true)
}
}(b) Types without defaults, built from decoded fields once all of them decoded: a builder that gathers every field's issues before calling the initialiser.
Encoding. .full writes every field; .sparse omits fields equal to
their default — the shape of a hand-edited config file. Canonical byte
output is StoicJSON's canonical encoding.
Secrets. A secret field never appears in actual, in description,
or in sparse output unless asked for.
Numbers. A JSON integer above 2^53 aimed at a Double field is an
issue (precision_loss), not a silent rounding. 1.0 for an Int field is
accepted only under .lenient. NaN and infinities do not exist in JSON and
cannot be produced.
Layer 3: tooling
- JSON Schema export, Draft 2020-12: ranges, lengths, patterns, enums,
defaults,
descriptionfromdoc,additionalProperties: falsefor rejecting objects,x-stoic-unitandx-stoic-secretmetadata. What cannot be expressed (cross-field constraints,.custom) is exported as anx-stoic-constraintannotation and listed in the export's summary. - Generators. Every schema derives valid values biased to boundaries
(lower, lower+1, upper−1, upper, empty and full collections, astral and
combining Unicode) and invalid documents — a valid one with one minimal
violation and the expected
(path, code)recorded. - SchemaSuite. A one-line self-test per schema: round trip, sparse
fixed point, every generated mutation yields exactly the expected issue at
the expected path, every typo'd key yields the right suggestion, and the
JSON Schema export agrees with
decodeon every generated document. - Migrations operate on
JSONValue, as a closed set of declarative operations (rename,move,delete,setIfAbsent,mapEnum,pinDefault,custom(id:)). Defaults are versioned: when a release changes a default, a document from the older version that omitted the field gets the old default written in explicitly by the migration, so nobody's chosen value flips silently. A golden document from every released version is kept and must still migrate to the latest. Documents newer than the app are read-only unless the schema preserves unknown fields.
Layer 4: sugar
@Schema on a struct with defaults emits form (a) from the stored
properties, their default values, attribute constraints (@Range,
@Length, @Secret, @Key) and doc comments; on an enumeration it emits
the enumeration or, when cases carry values, the tagged union with the two
closures per case; @Refined emits a named refinement. The macros live in a
separate StoicMacros product, behind the Macros trait, and only call the
public runtime (see "As built: Layer 4"). A Codable bridge lets a schema-backed type sit
inside Codable code: its init(from:) decodes the subtree through the
schema and throws one DecodingError listing every issue.
Failure modes
| What happens if… | Behaviour |
|---|---|
| three fields are invalid | Three issues, each with its path (and line/column from text); no value |
| a key is misspelt | unknown_key with "did you mean"; rejected, warned or ignored per policy |
| a key uses the snake_case spelling | Accepted as an alias; deprecated_key warning |
a number is out of range under .strict |
out_of_range error with the bounds |
the same under .lenient |
Default used; repaired issue saying so |
an integer exceeds 2^53 for a Double field |
precision_loss error |
| a secret field is invalid | Issue without actual |
| the input has duplicate keys | Rejected by StoicJSON with both locations |
| a default changed between releases | Old documents keep the value they had, via a pinDefault migration |
Testing
SchemaSuite runs on every schema in the test target. The decoder's accumulation is checked by the mutation generator: N independent mutations must produce exactly N issues at exactly their paths. The exporter is checked differentially against a small JSON Schema validator in the test target. Encoding is checked for round trips and sparse fixed points over generated values.
As built: phase 1 (StoicSchema)
Layers 1 and 2 and the JSON Schema exporter of Layer 3 are implemented in
StoicSchema, which depends on StoicJSON alone. The decisions the sketches
above left open:
- Constraints are a struct over a
Kindenum, soConstraint.range(…)reads as in the examples while exporters switch overconstraint.kind. Bounds areJSONNumbers and are compared as exact decimals (up to 38 digits), so2^53 + 1is above2^53and0.3is a multiple of0.1. A constraint is silent about kinds of value it is not about, as JSON Schema's keywords are. - String length has a unit.
.charactersis Unicode scalars, the only unit JSON Schema counts;.graphemesand.utf8Bytesexport asx-stoic-length. A pattern is skipped when a length constraint on the same value already failed. - Patterns are stored as source and compiled with
Regexper decode (Regexis notSendable), with scalar matching semantics. Refinementalso namesbaseSchema(automatic for anySchematicbase), because checking a value means interpreting the constraints on its JSON form.init(clamping:)exists for integer andDoublebases and clamps range constraints only.Optionalhas three states. Absent means the field's default (ornilwhere there is no default);nullis a value only fornullable: true; otherwisenilis written by omitting the key.- Object form (b) reads fields into
Slots and builds through aValuestoken, so every field's issues exist before the initialiser runs; see the note inSchema+ReadObject.swift. Form (a) defaults tounknownKeys: .warn, form (b) to.reject. - Respelled keys (case and snake_case, plus declared aliases) are
accepted with a
deprecated_keywarning; the same field under two spellings isduplicate_key. - Encoding.
EncodeStyle.sparseleaves out secrets unless asked;.fullincludes them. Sparse compares encoded JSON, not values. Deprecated fields are written only when they differ from their default. - Export.
jsonSchema()andjsonSchemaExport()(which also lists the rules that were only annotated). Recursive schemas use$defs/$ref; named non-recursive types are inlined.
As built: phase 2 (StoicSchema)
Phase 2 completes Layer 3 and the bridge of Layer 4. What was decided, in the order it was built:
Custom schemas
Schema.transform(base:decode:encode:describe:name:examples:)reads the base schema first and then converts, so the base's constraints hold (and all its issues are reported) before the closure runs. The closure throws aTransformFailure(message, code, expectation, suggestions), which becomes oneinvalid_valueissue at the value's path.encodeis total.- The tree describes it as
SchemaNode.transformed(TransformNode): the base node, the Swift type's name, a description andexamples. The export is the base's keywords plusx-stoic-transform {name, description}. The rule "constraints are data" survives in the only way it can for code: put everything the base can say in the base. - Examples exist for the generators. A generator cannot look inside a closure, so a transform declares encoded values it is known to accept; without them the generator tries base values and the suite says so.
Schema.custom(describing:description:decode:encode:)is the escape hatch for shapes no composition describes. Its decoder reports through aDecodingContext, which knows the path (context.report(…, at: ["low"])records$.range.low, with line and column when the input was text) and decodes sub-values with other schemas (context.decode(_, as:, at:)). A decoder that returnsnilwithout reporting still fails withinvalid_value; one that returns a value after reporting an error fails.Schema.enumeration(_:)also takesIntraw values (SchemaNode.integerEnumeration, with the Swift case names kept for documentation):type: integer,enum: [...],x-stoic-enum-names.
Tagged unions
Schema.union(discriminator:) { Case(…) … }declares an enumeration with payloads. EachCasenames its tag, its payload schema, how to build the enumeration from a payload (the case's own name, which is a function) and how to take a payload back out (a closure returningnilfor every other case). Swift has no case paths, so each case says it twice; the@Schemamacro writes the pair. A case without a payload isCase("point", Shape.point) { if case .point = $0 { true } else { false } }.- Flat or keyed. An object payload is written flat, beside the
discriminator (
{"type": "circle", "radius": 5}), which is how wire formats look and whatoneOf+constvalidates. Any other payload is written under a value key ({"type": "label", "value": "Hello"}). Alazypayload counts as not an object, since looking inside would recurse while the schema is built. - Decoding reads the discriminator first, because until it is known
nothing else can be: not an object is
type_mismatch; absent ismissing_keyat$.type; a non-string istype_mismatchthere; an unknown tag isnot_one_ofthere with the tags listed and the nearest suggested. The payload then decodes as any schema does, accumulating, and sees the object without its discriminator, so a payload that rejects unknown keys does not rejecttype. - Export is
oneOfwith one branch per case, each pinning the discriminator withconstand requiring it, plusx-stoic-discriminator. - Declaration errors trap: a tag twice, a flat payload that declares the
discriminator itself, a value key equal to the discriminator. A value that
is none of the declared cases is found by
validate(_:).
Versions and migrations
- Declared on the object, in both forms:
version: 4, migrations: […](plusversionKey:, default"version", andmissingVersion:). The version lives in a reserved key that is not a field (declaring a field with that key traps), is removed before the fields see it, and is written first on every encode, sparse included. The node records it asObjectNode.versioning. - Migrations are
Migration.from(1, to: 2, ops…)overJSONValue, with a closed set of operations as data:rename,move,delete,setIfAbsent,mapEnum,pinDefaultandcustom(id:). Paths are dot-separated keys; a path that is not in the document is skipped; a rename or move onto a taken key is amigration_failederror, never an overwrite.pinDefaultdoes whatsetIfAbsentdoes and says why: it writes the old default into a document that omitted the field, so that a release changing a default cannot flip what an old sparse file meant. - Decoding reads the version, applies the steps from it to the current
version one at a time, and records a
migratedwarning per step (from,to, the operations in words) before validating the migrated object, so errors are at the new paths. A version that is not a positive integer istype_mismatch/out_of_range; one no step reads isversion_unsupported; a step that cannot be applied ismigration_failed. A document with no version key means the oldest version (.oldest, the default, for a format that gained its key late), the current one (.current, for hand-written files) or an error. - Positions are coarse after a migration. A migrated object has been rewritten, so the text no longer shows what its issues are about: they point at the whole object, not at a line that may now mean something else. Documents that needed no migration keep exact positions.
- Newer documents. A version above the schema's is
version_too_new, an error, unless the object preserves unknown keys:unknownKeys: .preservewithpreservingUnknownKeysIn: \Type.extra(form a) orf.unknownKeys(\.extra)(form b) keeps every key the schema does not declare in aJSONObjectproperty and writes it back. The document is then read best-effort with aversion_too_newwarning, the newer version number is kept with the unknown keys, and encoding writes it back rather than the schema's own: an old release never downgrades or truncates a newer one's file. (A preserved typo is not noticed; prefer.warnfor a hand-edited file.) - Export states the current version (
const), makes the key required unless a missing key means current, and lists the history underx-stoic-version. The document describes the current format only: a validator rejects an old document that the decoder would migrate. - Time capsule. The pattern, in the test target, is a table of golden
documents, one per readable version, each with the value it meant, decoded
under today's schema; a second test fails when the history grows without a
capsule entry. See
PrefsCapsuleinVersionedFixtures.swift.
The Codable bridge
- Inbound.
schema.decode(from: decoder)reads whatever the decoder holds into aJSONValue(JSONValue(reading:): a keyed container walked with aDynamicKey: CodingKey, else an unkeyed one, else a single value, tried in that order), runs the schema, and returns the value or throws oneDecodingError.dataCorruptedwhose description lists every issue with its full path from the document root (the decoder's coding path, then the schema's) and whoseunderlyingErroris theValidationFailure. Accumulation survives inside the subtree; the surroundingCodablecode still stops at its first error, as it must. - Outbound.
schema.encode(value, to: encoder)writes the schema's JSON through the encoder (JSONValue.write(to:)): objects as keyed containers, arrays as unkeyed, integers asInt64/UInt64, other numbers asDouble. A value the schema cannot write (NaN) is anEncodingError. SchemaCoded<T: Schematic>is the wrapper (Codable,EquatableandHashablewhenTis) for a property, an element or a top-level value;decodeSchema/decodeSchemaIfPresent/encodeSchemaon the containers serve hand-writteninit(from:).CodingUserInfoKey.schemaDecodeOptionsand.schemaEncodeStylechoose lenient decoding or sparse encoding throughuserInfo, without touching any conformance.- What
Decodablecosts. ADecoderhides the source text: numbers arrive asInt64,UInt64orDouble(so1.0is1and a negative zero is0, and a non-integer is as precise as aDouble), object members arrive sorted by key, andJSONDecodercannot tell"é"from"e\u{301}"in a key. There are no source ranges. The schema's own text entry points keep all of that; use them for files and the bridge forCodablegraphs. - Not provided:
Schema.codable(T.self)(wrapping an existingCodabletype as a leaf) and theuserInfocollector that lets a generatedinit(from:)continue after a failed field (both are in the review's 3.9; neither is needed without macros). - Tested against a minimal
Decoder/EncoderoverJSONValuewritten in the test target with no Foundation, and, in the one test file that imports Foundation, againstJSONDecoder/JSONEncoderandPropertyListDecoder/PropertyListEncoder. A test also scans the sources ofStoicSchemaandStoicJSONforimport Foundation.
SchemaSuite
try SchemaSuite(Settings.schema).run(seed: 1, cases: 500) is the
one-line self-test. It generates valid documents and checks: round trip (as
JSON, canonical text and pretty text; validate agrees); sparse fixed point
(with and without secrets); that every minimal mutation yields exactly the
expected errors at exactly the expected paths and the expected warnings;
that two to four independent mutations yield exactly that many errors; that a
misspelt key or tag is answered with the spelling it came from; that the
exported JSON Schema and the decoder agree on the valid and the mutated
documents; that an empty document decodes to the declared defaults and the
defaults are valid; that every field has documentation (a warning); and the
migration chain.
- It lives in
StoicSchema, and asserts nothing.report(seed:cases:)returns aSchemaSuiteReport(a tally per check, the first findings of each with the seed that reproduces them and a shrunk document);runthrowsSchemaSuiteFailureif any check failed. A throw fails a Swift Testing or XCTest test and prints every finding, and the library apps ship never links a test framework. - The Swift Testing glue is a separate tiny target,
StoicSchemaTesting(expectSchemaSuite(_:seed:cases:), which records each finding as an issue at the caller's line, warnings as warnings). It is not inStoicTestingbecause that would make every user of virtual clocks link the schema layer. - The differential validator moved into the library as
JSONSchemaValidator(public, deliberately small, only the keywords the exporter writes), because the suite needs it. Rules a JSON Schema cannot state are skipped by name in the agreement check: precision loss and overflow of aDouble, and lengths counted in graphemes or bytes. - Failures are shrunk where the property is about a value (round trip), so the report shows the smallest document that still fails.
- A transform with no examples makes the generator guess, so the suite
says so (a warning) and skips samples the conversion refuses, counting them
as
skipped.
The migration chain check
schema.migrationChainCheck(cases:seed:) returns a MigrationChainReport.
The old schemas are gone (that is the point of migrating JSON), so it does
not generate old documents from them. It generates valid current documents,
runs the history backwards (the inverse of a rename is the rename back; of
a move, the move back; of pinDefault and setIfAbsent, a delete, so the old
document omits the field and the pin must supply it; of a mapEnum, the
reversed mapping if it is injective; of a delete, restoring its example;
of a custom, its inverse), and decodes the result with the full schema:
every old document must come out valid, with one migrated notice per step.
A step is untested, not passed, if something after it has no inverse, and a
delete with no example is reported as a gap. It also fails a step whose
rename or move writes a path no generated document of the next version has,
since such a step names a field the schema does not have and nothing in the
backwards documents would reveal it.
The check is the complement of the time capsule: the capsule proves the files that exist; the chain proves the ones that could.
Reading order
Issues decoded from text are returned, and printed by ValidationFailure,
sorted by position in the file (line, then column; stable, so issues at one
position keep the order they were found in). A schema declares fields in the
order that suits the type, and a report that jumps up and down the file is
one people stop reading. Issues without a position (a value that was never
text) keep the order found, which is declaration order within an object and
element order within an array; there is no better "path order" for them,
since sorting paths alphabetically would be worse than the declaration.
A missing key sorts with the object that lacks it.
Generators
schema.generatoris aSchemaGenerator<Value>derived from the node tree alone; it decodes what it makes, so aSample(the document and the value) is accepted by definition, and a schema it cannot satisfy raises aGenerationFailurethat says where, instead of returning a bad sample. Randomness is a seededSplitMix64(public), drawn from the caller's ownRandomNumberGeneratoror fromsample(seed:).- Valid values are candidate-and-filter. The generator proposes the edges
(bounds, the values beside them,
lower+1,upper-1, zero, negative zero, 2^53, the smallest and largestDouble, empty and full collections, strings of the shortest and longest length) mixed with uniform draws, and the real constraint interpreter keeps what satisfies every constraint. That is howcustomandmultipleOfand an unsupported pattern are met without the generator knowing them: the checker is the judge. - Strings are measured in the constraint's unit (scalars, graphemes or
UTF-8 bytes) from pieces whose costs in all three units are known: ASCII,
accented letters, CJK, astral characters, combining sequences, a skin-tone
modifier, a ZWJ family, a flag, and the ASCII a JSON writer must escape.
Several units at once get plain text, the only text that counts the same
in all. Patterns are sampled by a small regex sampler (literals, classes,
groups, alternation, quantifiers; anything else declines) whose output is
verified with
Regex. - Documents come sparse, dense and mixed, so defaults, omitted keys and explicit values all occur; versioned objects carry their current version. Recursion is bounded by depth and a per-document node budget (without the budget a union with eight recursive terms makes millions of nodes); past either, collections are as small as allowed and options are left out.
- Transforms are met by their
examples. A closure cannot be proposed against, so the generator draws from the encodings of the examples the schema declares;transformsWithoutExampleslists the ones that have none. - Shrinking is structural.
shrink(_:)proposes simpler documents, simplest first (numbers toward zero or the nearest bound, strings toward the minimum length, collections toward their minimum count, fields toward their defaults or absence) and keeps those the schema accepts;minimize(_:whileFailing:)greedily follows them to a local minimum. - Invalid documents.
mutations(of:)walks the document against the tree and lists every minimal violation with the issue it must cause (ExpectedIssue: path, code, severity, and for a misspelling the key or tag to suggest): below, above, wrong type, fractional,null, missing required key, unknown key, misspelt key, too long, too short, too many, too few, repeated element, not a multiple, pattern mismatch, bad enumeration case, bad union tag, misspelt tag, missing discriminator, lost precision, not finite, future version. Every candidate is run through the constraint interpreter first so that "exactly one issue" is a property of the mutation.mutation(of:combining:)applies several whose parts of the document do not overlap (a wrong container hides what is inside it, so a mutation's location is its whole subtree where that is true, and a mutation in auniquearray makes the array one unit). - Not generated: duplicate keys (the JSON layer refuses them before a
schema exists), a depth bomb (the parser's limit), and violations of
customconstraints (opaque).
As built: Layer 4 (StoicMacros)
@Schema and @Refined are implemented in the StoicMacros product, which
needs the Macros package trait. The rule of the layer is kept: a macro
expansion is calls to the public StoicSchema API and nothing else. There
is no runtime support type, no hidden protocol, no generated code a person
could not have written. The test target proves it by declaring types with the
macros, declaring the same schemas by hand, and requiring the two to be
indistinguishable (same node tree, same JSON Schema, same verdict, value and
issues on every generated and mutated document, same bytes from every
encoding style); SchemaSuite then runs on the macro's schemas.
The product and the trait
Macrosis a package trait, arranged exactly asLintis: the swift-syntax dependency stays declared, its products are conditioned on the trait, and the sources ofStoicMacrosPlugin(the compiler plugin) compile to a stub withoutSTOIC_MACROS. A consumer that does not enable the trait never fetches swift-syntax (checked with a scratch package and an empty SwiftPM cache: no checkout, noPackage.resolved, nothing in the cache but manifests).import StoicMacroswithout the trait is an empty module, and@Schemais then an unknown attribute, which is the honest answer.- Three targets.
StoicMacrosPlugin(.macro, swift-syntax, host only),StoicMacros(the attribute declarations; depends onStoicSchemaand the plugin) and two test targets:StoicMacrosTests(expansion text and diagnostics, with swift-syntax's generic macro test support recording into Swift Testing) andStoicMacrosUsageTests(behaviour). - The consumer imports
StoicSchemaandStoicMacrosand nothing else for the generated code. The expansion names the library through module selectors (StoicSchema::Schema,StoicSchema::Field), so a type of the consumer's own calledSchemaorFieldcannot change what it means.Secret<T>fields nameSecret, which the consumer's own property type already needs. Apublictype needspublic import StoicSchema, as any public API built on it does; the macro does not re-export. - Names. The attributes are named for what they say, and several are
also the names of types (
Rangein the standard library,Unitin Foundation,Secretin Stoic,SchemaandRefinedin StoicSchema). Macros and types are looked up apart in attribute position, so@Range(1...5) var r: Range<Int>works in one file with Foundation imported;NamingTestsis the proof.
@Schema on a struct
/// Application settings.
@Schema
struct Settings {
/// Pixels per second.
@Range(200...4000) @Unit("px/s") var scrollSpeed: Double = 1200
var theme: Theme = .system
@Key("api_token") @Secret @Length(...128) var apiToken: String = ""
@Length(1...253) @Count(0...8) @Unique var hosts: [String] = []
var network = Network()
@Aliases("legacy") @Deprecated("Use `theme`.") var legacyMode = false
@Ignored var cache: [String: Int] = [:]
}expands to static let schema = Schema.object(Settings.self, default: Settings(), doc: …) { Field(…) … }, one Field per stored property in
declaration order, and extension Settings: Schematic {} (added only if the
type does not conform already), so Settings.schema is what other schemas
write when they hold a Settings.
- Form (a) only. The default instance is
Settings(), so every property needs a default (an optionalvarhas one), and a type that declares initialisers must declareinit(). The macro does not read defaults to copy them (the instance is the single source of every default, which is the point of form (a)); it reads them to check them, below. Form (b), for types without defaults, is not emitted (see open problems). - Types are read as written. The macro sees syntax.
Bool,Int,Int64,UInt64,Double,String,JSONValueand the other fixed-width integers (.integer(Int32.self)) are the library's primitives;[T]and[String: T]are.array(of:)and.dictionary(of:);T?is.optional(@Nullableaddsnullable: true);Secret<T>is a.transformoverT's schema that reads and reveals, and makes the field secret; any other name, including a typealias such asPortand a generic such asRefined<IntRange<1, 9>>, isName.schema, and the compiler says if it has none. The type's own name, anywhere inside a collection or an optional, is.lazy { Name.schema }, so a recursive type works. A type with no schema that the macro can recognise (Set,Date,URL,UUID,Float, a tuple, a function, an existential, a dictionary with non-Stringkeys) is refused by name, with what to use instead. - A property without a type annotation is typed from a literal default
(
1isInt,1.5isDouble, a string, a boolean) or from aName(…)initialiser call; anything else is refused ("write the type"). - Doc comments are the fields'
doc:, and the type's is the object'sdoc:. Lines of a paragraph are joined by a space; a blank///line starts a new paragraph. @Schema's arguments are copied toSchema.objectunread:name,unknownKeys,version,versionKey,missingVersion,migrations.@PreserveUnknownKeysmarks theJSONObjectproperty that becomespreservingUnknownKeysIn:and is not a field.- Access.
publicandpackagetypes get apublic/packageschema, because the conformance needs a witness as visible as the type.
Constraint attributes
| Attribute | Is | Applies to |
|---|---|---|
@Range(200...4000) |
.range(…) |
numbers |
@GreaterThan(0), @LessThan(1) |
.greaterThan, .lessThan |
numbers |
@MultipleOf(5) |
.multipleOf |
numbers |
@Length(1...64, unit:) |
.length(…, unit:) |
strings |
@Pattern("^[a-z]+$") |
.pattern |
strings |
@Count(0...8) |
.count |
the outermost array or dictionary |
@Unique |
.unique |
the outermost array |
Each is written beside its property, in the order written. Scalar
constraints reach the innermost value through optionals, arrays,
dictionaries and secrets (@Length(1...253) var hosts: [String] bounds each
host; @Count(0...8) bounds the list), because that is the only reading that
lets both be said with one attribute each. A constraint cannot reach inside a
named type the macro cannot see (a typealias, a struct); that is an error that
says to put it in the type's own schema.
The other markers: @Key("wire"), @Secret, @Unit("px/s"),
@Aliases("a", "b"), @Deprecated("why"), @Nullable, @Ignored (the
property stays in the default instance and out of the schema),
@PreserveUnknownKeys. All of them are peer macros that expand to nothing:
@Schema reads them from the syntax of the members it is given, and each one
used outside a @Schema type is an error, not silence.
@Schema on an enumeration
- A
StringorIntraw type givesSchema.enumeration(T.self), and theCaseIterableconformance it needs if the enumeration lacks it (the extension macro adds only what is missing). - Cases that carry values give
Schema.union(T.self, discriminator:), with the pair of closures per case that phase 2 had each case say by hand:Case("circle", Circle.schema, Shape.circle) { if case .circle(let payload) = $0 { payload } else { nil } }. A case with no values isCase("point", Shape.point) { if case .point = $0 { true } else { false } }. One associated value, labelled or not, is the payload; its schema is the type's (writtenSchema<T>.…so the compiler has a type to infer from). Several labelled values are an object of their own, read with the form (b) reader over a labelled tuple,Shape.rect(width: 3, height: 4)being{"type": "rect", "width": 3, "height": 4}.@Key("pt")on a case renames its tag.@Schema(discriminator:, valueKey:, name:)are passed on. - Constraints on payloads have nowhere to be written (Swift has no attributes on associated values), so a payload's constraints live in its type's schema.
- A plain enumeration with neither raw type nor values is refused, with a
fix-it that adds
: String.
@Refined
@Refined(Double.self, .range(0.1...10), .finite) enum ScrollSpeedRule {}
adds typealias Base, static var constraints and the Refinement
conformance. The constraints are the library's own Constraint values, copied
unchanged; they are checked with the same code as the attributes (does each fit
the base, is a range empty, do they contradict), but, as in a refinement, they
are about the base value itself: a number constraint does not reach inside
[Int].self. The value type stays Refined<ScrollSpeedRule>.
Diagnostics
A macro that finds an error reports it at the node that is wrong, with a
fix-it where one is safe, and expands to nothing (no schema, no
conformance), so the person sees the cause and not a cascade of "type has no
member schema". Every diagnostic has a test with its line and column.
| Where | Mistake |
|---|---|
| the type | not a struct or enum; generic; an enum with no cases; already declares schema; initialisers but no init() |
| a property | no default (fix-it adds 0, "", false, [], [:]); a let (fix-it: var); type not visible (no annotation, non-literal default); type with no schema, by name; a pattern binding (var (a, b)) |
| use of markers | any marker on a static or computed property, on an @Ignored property, or outside a @Schema type; the same marker twice (fix-it removes the second); @Key(""); a non-literal key, alias or deprecation message |
| keys | a key or alias used twice (the runtime precondition, moved to compile time); an alias equal to its own key; the version key used by a field |
| constraints | one that does not fit the type (@Length on Int, @Range on String or on a named type, @Count on a scalar, @Unique on a dictionary; fix-it removes it); an empty range (10...1); a negative length or count; @MultipleOf(0); constraints that exclude each other; a pattern that does not compile; a literal default that breaks its own constraint |
| nullable and secret | @Nullable on a non-optional (fix-it removes); @Secret on a Secret<T> (warning) |
| arguments | discriminator: or valueKey: on a struct, unknownKeys: or version: on an enum (fix-it removes); migrations:, versionKey: or missingVersion: without version:; a version below 1; unknownKeys: together with @PreserveUnknownKeys; .preserve with no preserving property; a preserving property that is not a JSONObject, or a second one |
| an enum | no raw type and no values (fix-it adds : String); a raw type other than String or Int; a marker on a raw-valued case; a marker other than @Key on a union case; a tag used twice; the discriminator equal to the value key; several associated values without labels; an associated value of a type with no schema |
@Refined |
on anything but an enum or struct; no base type or constraints; a base that is not T.self; a constraint that does not fit the base, an empty range, contradictions |
A non-literal default is not an error: the default instance carries it, and
the macro checks only what it can see (a literal number against @Range, a
literal string against @Length, a literal array against @Count and
@Unique). What it cannot see, SchemaSuite checks ("the defaults are
valid").
Decisions
- Sugar means the tests compare against hand-written schemas, not against expected text alone. The expansion text is also tested, for review.
Secret<T>is a transform, not a new schema in StoicSchema, because StoicSchema does not depend on Stoic andSecretlives there.- Self-reference is lazy, other references are direct. A
.lazypayload of a union counts as "not an object" (phase 2), which would change the wire shape of every nested payload, so only the type's own name is lazy. - Checks are literal arithmetic. The macro evaluates integer, floating
point, string, array and dictionary literals and range expressions, and
passes everything else through, so
@Range(Self.limits)is legal and unchecked until the schema is built.
Cost
On 40 structs of 8 fields each (320 fields, constraints on half of them), building the module with the macros and with the same schemas written by hand took the same time within noise (3.2–4.2 s against 2.9–4.9 s per debug build of the target, including the link), so expansion is about ten milliseconds per type. The one-time cost is building the plugin: a fresh consumer package (StoicJSON, StoicSchema, StoicMacros, the plugin, and swift-syntax's prebuilt macro support, 0.7 s to download) built in 7 s. Without the prebuilt (an older toolchain, a locked-down machine) SwiftPM builds swift-syntax from source, which takes minutes; that is the price of the trait and the reason it is a trait.
Open problems
- Form (b) is not emitted. A struct without defaults needs the reader
over the memberwise initialiser; the macro refuses it instead. The
generated
build { v in T(a: v[a], …) }is mechanical, but the memberwise initialiser's labels and access level are not visible from the declaration when it is not written out. - Mutual recursion (
Aholds[B],BholdsA?) is not detected; bothstatic lets then wait for each other and trap on first use. Only a type's own name is made lazy. A@Recursivemarker that wraps a field in.lazyis the likely fix. - Type aliases and other opaque names are
Name.schema. A constraint on one is refused rather than guessed. - Documentation of enumeration cases is not exported:
Schema.enumerationhas nowhere to put it (union cases do takedoc:). - A union's recursive payload is keyed under the value key (
.lazyis not an object), which differs from its non-recursive siblings' flat form; this is phase 2's rule showing through. - A non-literal constraint argument is unchecked at compile time and traps (the library's own precondition) when the schema is first built.
- Property wrappers,
lazyproperties and observed macros such as@Observableare not understood; a macro that rewrites stored properties before@Schemasees them is outside what it reads. - Linux and Windows are outside the platform floor, as for the rest of Stoic; the plugin itself has no platform dependency.
Non-goals
Formats other than JSON in the first version (property lists and YAML can
map onto JSONValue later); remote schema registries; validation that needs
I/O (it belongs in an async second stage that reports the same Issue type).