The schema term: what it is, how to check one, and how to rewrite one.
A schema is a tuple, {kind, opts} or {kind, payload, opts}, built by Rupa.T. Nothing
in it is a function, so a schema prints, compares, hashes with :erlang.phash2/1, and
embeds in generated code as a literal. validate/1 is the only thing that judges a schema;
the constructors themselves hand back whatever you gave them.
Present, absent, null
Three states, kept apart on purpose.
- A field is required unless its term is wrapped in
Rupa.T.optional/2or carriesdefault:. A required key must be present and must not benull. Rupa.T.optional/2says the key may be absent. Absent decodes to the key simply not being in the map — not tonil.Rupa.T.nullable/2says the value may be JSON null, which decodes tonil. It says nothing about the key being absent.default:fires on absent only. An explicitnullis never replaced by a default; it is either allowed, bynullable, or an error.
So T.optional(T.nullable(T.string())) is the field that may be missing, may be null, and is
otherwise a string — and decoding tells you which of the three you got.
Field names
A field's name is an atom or a string, and an object's names are all of one or all of the
other — mixing them is :mixed_field_names, because it would make the decoded map's shape
depend on which field you were looking at.
Write atoms when you know the fields as you write the code, which is most of the time. Strings
are for an object whose field names you did not choose — one that came out of
Rupa.JsonSchema.decode/1, say — and they are what makes that path total: nothing is interned
from a document that arrived over the wire.
Names on the wire
A field's wire key is a string, and by default it is the name spelled out. Four options move that apart, all of them settled at compile time.
rename_all:is an object option —:camelCase,:snake_caseor:kebab. It reads each field name as words and writes them back in that style, both directions, so%{first_name: ...}withrename_all: :camelCasereads and writes"firstName". It applies to that object and not to the objects inside it;walk/2is how you apply one to a whole tree.from:andto:override it for one field:from:is the key the value arrives under,to:the key it leaves under. Give one and the other follows it, so a field renamed by hand still round-trips; give both, naming different keys, and you have said that this field reads an old name and writes a new one — the one way to makeRupa.encode/3stop beingRupa.decode/3's inverse, which is exactly what a migration wants. Likedefault:, they belong on the field's own term — the node in the object's field map,Rupa.T.optional/2included — and anywhere else is:misplaced_rename.keys:is the object option for the other side::atomor:string, deciding the type of the key the decoded map holds. It changes the key's type, never its name, so it stays out of the wire's way — and encoding reads back whichever it wrote. It defaults to the type of the name you wrote, so an atom-named object decodes to atom keys and a string-named one to string keys, and neither default interns anything.
Renaming never interns an atom from wire data: the atoms are the field names, and they exist
before any data does. keys: :atom on a string-named object is the one place a name becomes
an atom, and it is a name in a schema rather than a key off a document — the same rule every
other atom in a schema already follows. Two fields that land on one wire key are refused at
compile time rather than one of them quietly winning.
Structs
into: on an object names a struct module to decode into. It is an option on the object, not
on the compile, so nesting works: an address three levels down becomes a MyApp.Address the
same way the root becomes a MyApp.User. A module name is an atom, so the schema still
prints, hashes and embeds as a literal; what it adds is a soft dependency on that module,
which Rupa.compile/2 checks — the module must be loaded and must already have every key the
object decodes to, or you get a schema error rather than a surprise at decode time.
mix rupa.gen.struct writes those modules from the schema.
A struct has atom keys and every one of them always, which rules three things out. keys:
must resolve to :atom, so keys: :string is refused and a string-named object needs the
explicit keys: :atom that interning already makes you ask for. unknown: :keep is refused,
because a struct has nowhere to put the keys it would carry. And Rupa.T.optional/2 keeps
working but stops meaning three things: the key is always there, so what decoding leaves
behind for an absent field is that field's defstruct default, and encoding writes that same
value back as absent — which is what keeps the round trip closing. For a struct
mix rupa.gen.struct wrote, that value is nil, so an absent key and an explicit null reach
the same struct and both leave as absent. A struct you wrote yourself may default the field to
something else, and then that something else is what stands for absent, and nil is a value
like any other.
Paths
Errors from validate/1 carry a path through the schema:
- an atom is an object field, or a
defs:entry after the:defssegment :ofis a list's element type or amap_of's value type- an integer is a position in a tuple or a variant of a union
- a string is a branch of a tagged union, or a field of a string-named object
Summary
Functions
The built-in formats, in the order they are documented.
Puts a schema in canonical form: bare maps become objects, and options are sorted by key.
Checks that a term is a well-formed schema, and returns it in canonical form.
validate/1, raising Rupa.SchemaError instead of returning errors.
Rewrites every node of a schema, children first.
Types
@type name() :: atom()
@type opts() :: keyword()
@type t() :: {:string | :integer | :float | :boolean | :null, opts()} | {:literal, term(), opts()} | {:enum, [term()], opts()} | {:ref, name(), opts()} | {:object, %{required(atom() | String.t()) => t()}, opts()} | {:list | :map_of | :optional | :nullable, t(), opts()} | {:tuple | :union, [t()], opts()} | {:tagged, %{required(String.t()) => t()}, opts()}
Functions
@spec formats() :: [atom()]
The built-in formats, in the order they are documented.
iex> :uuid in Rupa.Schema.formats()
true
Puts a schema in canonical form: bare maps become objects, and options are sorted by key.
Rupa.T does this as it builds, so a term you did not hand-write is already canonical. Two
schemas that mean the same thing are then the same term, which is what makes
:erlang.phash2/1 a usable identity for a compiled codec.
iex> Rupa.Schema.normalize(%{name: {:string, []}})
{:object, %{name: {:string, []}}, []}
iex> Rupa.Schema.normalize({:string, [min: 1, len: 2]})
{:string, [len: 2, min: 1]}
@spec validate(term()) :: {:ok, t()} | {:error, [Rupa.Error.t()]}
Checks that a term is a well-formed schema, and returns it in canonical form.
This is structure only: that every node is a kind Rupa knows, that its options are ones that
kind takes and hold values of the right shape, that no branch of the tree holds a function,
and that every Rupa.T.ref/2 resolves. It says nothing about whether any data would decode.
Every problem is reported, not just the first.
iex> Rupa.Schema.validate(%{name: {:string, [min: 1]}})
{:ok, {:object, %{name: {:string, [min: 1]}}, []}}
iex> {:error, [error]} = Rupa.Schema.validate({:string, [format: :ssn]})
iex> {error.path, error.code}
{[], :unknown_format}
validate/1, raising Rupa.SchemaError instead of returning errors.
iex> Rupa.Schema.validate!({:boolean, []})
{:boolean, []}
Rewrites every node of a schema, children first.
The function is handed each node after its children have been rewritten, and whatever it
returns takes that node's place. Nodes inside defs: are rewritten too.
The schema is normalised first, so the function always sees canonical {kind, ...} nodes: a
bare map arrives as {:object, fields, opts} -- and is descended into -- rather than as the
leaf it would otherwise look like. That is what lets one {:object, _, _} clause restyle every
object in a tree written the idiomatic bare-map way.
iex> schema = %{age: {:integer, []}}
iex> Rupa.Schema.walk(schema, fn
...> {:integer, opts} -> {:integer, [{:gte, 0} | opts]}
...> node -> node
...> end)
{:object, %{age: {:integer, [gte: 0]}}, []}