Rupa.Schema (rupa v0.1.0)

Copy Markdown View Source

The schema term: what it is, how to check one, and how to rewrite one.

A schema is a tuple, {kind, opts} or {kind, payload, opts}, built by Rupa.T. Nothing in it is a function, so a schema prints, compares, hashes with :erlang.phash2/1, and embeds in generated code as a literal. validate/1 is the only thing that judges a schema; the constructors themselves hand back whatever you gave them.

Present, absent, null

Three states, kept apart on purpose.

  • A field is required unless its term is wrapped in Rupa.T.optional/2 or carries default:. A required key must be present and must not be null.
  • Rupa.T.optional/2 says the key may be absent. Absent decodes to the key simply not being in the map — not to nil.
  • Rupa.T.nullable/2 says the value may be JSON null, which decodes to nil. It says nothing about the key being absent.
  • default: fires on absent only. An explicit null is never replaced by a default; it is either allowed, by nullable, or an error.

So T.optional(T.nullable(T.string())) is the field that may be missing, may be null, and is otherwise a string — and decoding tells you which of the three you got.

Field names

A field's name is an atom or a string, and an object's names are all of one or all of the other — mixing them is :mixed_field_names, because it would make the decoded map's shape depend on which field you were looking at.

Write atoms when you know the fields as you write the code, which is most of the time. Strings are for an object whose field names you did not choose — one that came out of Rupa.JsonSchema.decode/1, say — and they are what makes that path total: nothing is interned from a document that arrived over the wire.

Names on the wire

A field's wire key is a string, and by default it is the name spelled out. Four options move that apart, all of them settled at compile time.

  • rename_all: is an object option — :camelCase, :snake_case or :kebab. It reads each field name as words and writes them back in that style, both directions, so %{first_name: ...} with rename_all: :camelCase reads and writes "firstName". It applies to that object and not to the objects inside it; walk/2 is how you apply one to a whole tree.
  • from: and to: override it for one field: from: is the key the value arrives under, to: the key it leaves under. Give one and the other follows it, so a field renamed by hand still round-trips; give both, naming different keys, and you have said that this field reads an old name and writes a new one — the one way to make Rupa.encode/3 stop being Rupa.decode/3's inverse, which is exactly what a migration wants. Like default:, they belong on the field's own term — the node in the object's field map, Rupa.T.optional/2 included — and anywhere else is :misplaced_rename.
  • keys: is the object option for the other side: :atom or :string, deciding the type of the key the decoded map holds. It changes the key's type, never its name, so it stays out of the wire's way — and encoding reads back whichever it wrote. It defaults to the type of the name you wrote, so an atom-named object decodes to atom keys and a string-named one to string keys, and neither default interns anything.

Renaming never interns an atom from wire data: the atoms are the field names, and they exist before any data does. keys: :atom on a string-named object is the one place a name becomes an atom, and it is a name in a schema rather than a key off a document — the same rule every other atom in a schema already follows. Two fields that land on one wire key are refused at compile time rather than one of them quietly winning.

Structs

into: on an object names a struct module to decode into. It is an option on the object, not on the compile, so nesting works: an address three levels down becomes a MyApp.Address the same way the root becomes a MyApp.User. A module name is an atom, so the schema still prints, hashes and embeds as a literal; what it adds is a soft dependency on that module, which Rupa.compile/2 checks — the module must be loaded and must already have every key the object decodes to, or you get a schema error rather than a surprise at decode time. mix rupa.gen.struct writes those modules from the schema.

A struct has atom keys and every one of them always, which rules three things out. keys: must resolve to :atom, so keys: :string is refused and a string-named object needs the explicit keys: :atom that interning already makes you ask for. unknown: :keep is refused, because a struct has nowhere to put the keys it would carry. And Rupa.T.optional/2 keeps working but stops meaning three things: the key is always there, so what decoding leaves behind for an absent field is that field's defstruct default, and encoding writes that same value back as absent — which is what keeps the round trip closing. For a struct mix rupa.gen.struct wrote, that value is nil, so an absent key and an explicit null reach the same struct and both leave as absent. A struct you wrote yourself may default the field to something else, and then that something else is what stands for absent, and nil is a value like any other.

Paths

Errors from validate/1 carry a path through the schema:

  • an atom is an object field, or a defs: entry after the :defs segment
  • :of is a list's element type or a map_of's value type
  • an integer is a position in a tuple or a variant of a union
  • a string is a branch of a tagged union, or a field of a string-named object

Summary

Functions

The built-in formats, in the order they are documented.

Puts a schema in canonical form: bare maps become objects, and options are sorted by key.

Checks that a term is a well-formed schema, and returns it in canonical form.

validate/1, raising Rupa.SchemaError instead of returning errors.

Rewrites every node of a schema, children first.

Types

name()

@type name() :: atom()

opts()

@type opts() :: keyword()

t()

@type t() ::
  {:string | :integer | :float | :boolean | :null, opts()}
  | {:literal, term(), opts()}
  | {:enum, [term()], opts()}
  | {:ref, name(), opts()}
  | {:object, %{required(atom() | String.t()) => t()}, opts()}
  | {:list | :map_of | :optional | :nullable, t(), opts()}
  | {:tuple | :union, [t()], opts()}
  | {:tagged, %{required(String.t()) => t()}, opts()}

Functions

formats()

@spec formats() :: [atom()]

The built-in formats, in the order they are documented.

iex> :uuid in Rupa.Schema.formats()
true

normalize(schema)

@spec normalize(term()) :: term()

Puts a schema in canonical form: bare maps become objects, and options are sorted by key.

Rupa.T does this as it builds, so a term you did not hand-write is already canonical. Two schemas that mean the same thing are then the same term, which is what makes :erlang.phash2/1 a usable identity for a compiled codec.

iex> Rupa.Schema.normalize(%{name: {:string, []}})
{:object, %{name: {:string, []}}, []}

iex> Rupa.Schema.normalize({:string, [min: 1, len: 2]})
{:string, [len: 2, min: 1]}

validate(schema)

@spec validate(term()) :: {:ok, t()} | {:error, [Rupa.Error.t()]}

Checks that a term is a well-formed schema, and returns it in canonical form.

This is structure only: that every node is a kind Rupa knows, that its options are ones that kind takes and hold values of the right shape, that no branch of the tree holds a function, and that every Rupa.T.ref/2 resolves. It says nothing about whether any data would decode.

Every problem is reported, not just the first.

iex> Rupa.Schema.validate(%{name: {:string, [min: 1]}})
{:ok, {:object, %{name: {:string, [min: 1]}}, []}}

iex> {:error, [error]} = Rupa.Schema.validate({:string, [format: :ssn]})
iex> {error.path, error.code}
{[], :unknown_format}

validate!(schema)

@spec validate!(term()) :: t()

validate/1, raising Rupa.SchemaError instead of returning errors.

iex> Rupa.Schema.validate!({:boolean, []})
{:boolean, []}

walk(schema, fun)

@spec walk(t(), (t() -> t())) :: t()

Rewrites every node of a schema, children first.

The function is handed each node after its children have been rewritten, and whatever it returns takes that node's place. Nodes inside defs: are rewritten too.

The schema is normalised first, so the function always sees canonical {kind, ...} nodes: a bare map arrives as {:object, fields, opts} -- and is descended into -- rather than as the leaf it would otherwise look like. That is what lets one {:object, _, _} clause restyle every object in a tree written the idiomatic bare-map way.

iex> schema = %{age: {:integer, []}}
iex> Rupa.Schema.walk(schema, fn
...>   {:integer, opts} -> {:integer, [{:gte, 0} | opts]}
...>   node -> node
...> end)
{:object, %{age: {:integer, [gte: 0]}}, []}