# `Rupa.Schema`
[🔗](https://github.com/zero-one-group/rupa/blob/v0.1.0/lib/rupa/schema.ex#L1)

The schema term: what it is, how to check one, and how to rewrite one.

A schema is a tuple, `{kind, opts}` or `{kind, payload, opts}`, built by `Rupa.T`. Nothing
in it is a function, so a schema prints, compares, hashes with `:erlang.phash2/1`, and
embeds in generated code as a literal. `validate/1` is the only thing that judges a schema;
the constructors themselves hand back whatever you gave them.

## Present, absent, null

Three states, kept apart on purpose.

  * A field is **required** unless its term is wrapped in `Rupa.T.optional/2` or carries
    `default:`. A required key must be present and must not be `null`.
  * `Rupa.T.optional/2` says the **key may be absent**. Absent decodes to the key simply not
    being in the map — not to `nil`.
  * `Rupa.T.nullable/2` says the **value may be JSON null**, which decodes to `nil`. It says
    nothing about the key being absent.
  * `default:` fires on absent only. An explicit `null` is never replaced by a default; it is
    either allowed, by `nullable`, or an error.

So `T.optional(T.nullable(T.string()))` is the field that may be missing, may be null, and is
otherwise a string — and decoding tells you which of the three you got.

## Field names

A field's name is an atom or a string, and an object's names are all of one or all of the
other — mixing them is `:mixed_field_names`, because it would make the decoded map's shape
depend on which field you were looking at.

Write atoms when you know the fields as you write the code, which is most of the time. Strings
are for an object whose field names you did not choose — one that came out of
`Rupa.JsonSchema.decode/1`, say — and they are what makes that path total: nothing is interned
from a document that arrived over the wire.

## Names on the wire

A field's wire key is a string, and by default it is the name spelled out. Four options move
that apart, all of them settled at compile time.

  * `rename_all:` is an **object** option — `:camelCase`, `:snake_case` or `:kebab`. It reads
    each field name as words and writes them back in that style, both directions, so
    `%{first_name: ...}` with `rename_all: :camelCase` reads and writes `"firstName"`. It
    applies to that object and not to the objects inside it; `walk/2` is how you apply one to
    a whole tree.
  * `from:` and `to:` override it for one field: `from:` is the key the value arrives under,
    `to:` the key it leaves under. Give one and the other follows it, so a field renamed by
    hand still round-trips; give both, naming different keys, and you have said that this
    field reads an old name and writes a new one — the one way to make `Rupa.encode/3` stop
    being `Rupa.decode/3`'s inverse, which is exactly what a migration wants. Like `default:`,
    they belong on the field's own term — the node in the object's field map,
    `Rupa.T.optional/2` included — and anywhere else is `:misplaced_rename`.
  * `keys:` is the **object** option for the other side: `:atom` or `:string`, deciding the
    type of the key the decoded map holds. It changes the key's type, never its name, so it
    stays out of the wire's way — and encoding reads back whichever it wrote. It defaults to
    the type of the name you wrote, so an atom-named object decodes to atom keys and a
    string-named one to string keys, and neither default interns anything.

Renaming never interns an atom from wire data: the atoms are the field names, and they exist
before any data does. `keys: :atom` on a string-named object is the one place a name becomes
an atom, and it is a name in a schema rather than a key off a document — the same rule every
other atom in a schema already follows. Two fields that land on one wire key are refused at
compile time rather than one of them quietly winning.

## Structs

`into:` on an object names a struct module to decode into. It is an option on the object, not
on the compile, so nesting works: an address three levels down becomes a `MyApp.Address` the
same way the root becomes a `MyApp.User`. A module name is an atom, so the schema still
prints, hashes and embeds as a literal; what it adds is a soft dependency on that module,
which `Rupa.compile/2` checks — the module must be loaded and must already have every key the
object decodes to, or you get a schema error rather than a surprise at decode time.
`mix rupa.gen.struct` writes those modules from the schema.

A struct has atom keys and every one of them always, which rules three things out. `keys:`
must resolve to `:atom`, so `keys: :string` is refused and a string-named object needs the
explicit `keys: :atom` that interning already makes you ask for. `unknown: :keep` is refused,
because a struct has nowhere to put the keys it would carry. And `Rupa.T.optional/2` keeps
working but stops meaning three things: the key is always there, so what decoding leaves
behind for an absent field is that field's `defstruct` default, and encoding writes that same
value back as absent — which is what keeps the round trip closing. For a struct
`mix rupa.gen.struct` wrote, that value is `nil`, so an absent key and an explicit null reach
the same struct and both leave as absent. A struct you wrote yourself may default the field to
something else, and then that something else is what stands for absent, and `nil` is a value
like any other.

## Paths

Errors from `validate/1` carry a path through the schema:

  * an atom is an object field, or a `defs:` entry after the `:defs` segment
  * `:of` is a list's element type or a `map_of`'s value type
  * an integer is a position in a tuple or a variant of a union
  * a string is a branch of a tagged union, or a field of a string-named object

# `name`

```elixir
@type name() :: atom()
```

# `opts`

```elixir
@type opts() :: keyword()
```

# `t`

```elixir
@type t() ::
  {:string | :integer | :float | :boolean | :null, opts()}
  | {:literal, term(), opts()}
  | {:enum, [term()], opts()}
  | {:ref, name(), opts()}
  | {:object, %{required(atom() | String.t()) =&gt; t()}, opts()}
  | {:list | :map_of | :optional | :nullable, t(), opts()}
  | {:tuple | :union, [t()], opts()}
  | {:tagged, %{required(String.t()) =&gt; t()}, opts()}
```

# `formats`

```elixir
@spec formats() :: [atom()]
```

The built-in formats, in the order they are documented.

    iex> :uuid in Rupa.Schema.formats()
    true

# `normalize`

```elixir
@spec normalize(term()) :: term()
```

Puts a schema in canonical form: bare maps become objects, and options are sorted by key.

`Rupa.T` does this as it builds, so a term you did not hand-write is already canonical. Two
schemas that mean the same thing are then the same term, which is what makes
`:erlang.phash2/1` a usable identity for a compiled codec.

    iex> Rupa.Schema.normalize(%{name: {:string, []}})
    {:object, %{name: {:string, []}}, []}

    iex> Rupa.Schema.normalize({:string, [min: 1, len: 2]})
    {:string, [len: 2, min: 1]}

# `validate`

```elixir
@spec validate(term()) :: {:ok, t()} | {:error, [Rupa.Error.t()]}
```

Checks that a term is a well-formed schema, and returns it in canonical form.

This is structure only: that every node is a kind Rupa knows, that its options are ones that
kind takes and hold values of the right shape, that no branch of the tree holds a function,
and that every `Rupa.T.ref/2` resolves. It says nothing about whether any data would decode.

Every problem is reported, not just the first.

    iex> Rupa.Schema.validate(%{name: {:string, [min: 1]}})
    {:ok, {:object, %{name: {:string, [min: 1]}}, []}}

    iex> {:error, [error]} = Rupa.Schema.validate({:string, [format: :ssn]})
    iex> {error.path, error.code}
    {[], :unknown_format}

# `validate!`

```elixir
@spec validate!(term()) :: t()
```

`validate/1`, raising `Rupa.SchemaError` instead of returning errors.

    iex> Rupa.Schema.validate!({:boolean, []})
    {:boolean, []}

# `walk`

```elixir
@spec walk(t(), (t() -&gt; t())) :: t()
```

Rewrites every node of a schema, children first.

The function is handed each node after its children have been rewritten, and whatever it
returns takes that node's place. Nodes inside `defs:` are rewritten too.

The schema is normalised first, so the function always sees canonical `{kind, ...}` nodes: a
bare map arrives as `{:object, fields, opts}` -- and is descended into -- rather than as the
leaf it would otherwise look like. That is what lets one `{:object, _, _}` clause restyle every
object in a tree written the idiomatic bare-map way.

    iex> schema = %{age: {:integer, []}}
    iex> Rupa.Schema.walk(schema, fn
    ...>   {:integer, opts} -> {:integer, [{:gte, 0} | opts]}
    ...>   node -> node
    ...> end)
    {:object, %{age: {:integer, [gte: 0]}}, []}

---

*Consult [api-reference.md](api-reference.md) for complete listing*
