public
When Validation Isn't Enough: Building Sch Library for OCaml
What started as an OpenAPI problem in Tapak became a library for defining schemas once and interpreting them in multiple ways.
Syaiful Bahri5 min readocaml

While working on Tapak’s OpenAPI support, I ran into a problem with validation. A validator can tell me whether an input is valid or not, but it doesn’t necessarily tell me what the input is supposed to look like.
That’s a problem when the same validation rules need to produce machine-readable documentation. OpenAPI and protocols such as MCP rely on structured descriptions of the data a service accepts.
I didn’t want to write those descriptions separately and keep them synchronized with the validation code.
Unfortunately, I couldn’t find an OCaml library that met these requirements. I needed type-safe validation, serialization, and an inspectable schema that could be translated into JSON Schema.
Many of the validation libraries I found use function combinators, with each rule represented as a function.
A rule might have the type string option -> ('a, string list) result, but no other interpreter can look inside that
function to produce JSON Schema.
That idea became Sch, an OCaml library for defining and validating schemas. You define a schema once and use it for validation, JSON encoding and decoding, or JSON Schema generation. JSON support is built in, but the schema layer is format-independent.
I wrote Sch for Tapak’s validation and OpenAPI generation, though it can be used outside Tapak.
The Problem in Tapak
The need for Sch emerged directly from Tapak’s handler architecture. Like Rust’s warp and
axum, Tapak does not restrict every handler to a request -> response pipeline.
A handler can receive values of any type, and the user defines how Tapak extracts them from the request.
Tapak must inspect those extraction rules to describe each route and generate the right OpenAPI document. One definition therefore has to extract and validate a value at runtime while also describing it in OpenAPI. Sch provides that definition.
Consider a route that accepts an age parameter constrained to a minimum of 18. Tapak needs to validate that parameter when handling a request, but it also needs to describe the same constraint in OpenAPI. If the validation rule is just a function, Tapak can execute it, but it cannot discover that the minimum age is 18.
Sch’s Design
Sch borrows its record design from Jsont and Caqti. Both build a record from a constructor function and project each field back out for encoding.
For example, Jsont uses pipeline syntax:
let jsont =
Jsont.Object.map ~kind:"User" (fun name email age -> { name; email; age })
|> Jsont.Object.mem ~enc:(fun p -> p.name) "name" Jsont.string
|> Jsont.Object.mem ~enc:(fun p -> p.email) "email" Jsont.string
|> Jsont.Object.mem ~enc:(fun p -> p.age) "age" Jsont.int
|> Jsont.Object.finish
I like this approach because the same definition supports both encoding and decoding. However, the constructor must accept each field as a positional argument, and the fields must be declared in that same order. This becomes cumbersome for records with many fields.
Sch keeps the projection model but uses applicative syntax instead:
let schema =
Sch.Object.(
define
@@ let+ name = mem ~enc:(fun u -> u.name) "name" Sch.string
and+ email = mem ~enc:(fun u -> u.email) "email" Sch.string
and+ age = mem ~enc:(fun u -> u.age) "age" Sch.int
in
{ name; email; age })
With let+ and and+, I can bind each field by name and construct the record using ordinary OCaml syntax, without
writing a separate positional constructor.
Internally, Sch uses nested tuples to combine these fields. The resulting schema can be interpreted for validation, serialization, deserialization, and JSON Schema generation.
Tapak also uses these definitions to deserialize query strings and multipart forms. The schema itself is independent of those formats, and each format provides its own interpretation.
One Schema, Several Uses
So far the schema only describes the shape of a record. Constraints are also part of the schema, so they can be inspected like everything else:
let schema =
Sch.Object.(
define ~kind:"User"
@@ let+ name =
mem ~enc:(fun u -> u.name) "name"
Sch.(with_ ~constraint_:(Constraint.min_length 2) string)
and+ email =
mem ~enc:(fun u -> u.email) "email"
Sch.(with_ ~constraint_:(Constraint.format `Email) string)
and+ age =
mem ~enc:(fun u -> u.age) "age"
Sch.(with_ ~constraint_:(Constraint.int_range 17 120) int)
in
{ name; email; age })
Decoding invalid input reports every failing field, not just the first one:
Sch.Json.decode_string schema {|{"name":"A","email":"nope","age":15}|}
|> Sch.Validation.to_result
(* Error
[ "age", "Integer 15 is less than minimum 17"
; "email", "Invalid email format"
; "name", "String length 1 is less than minimum 2" ] *)
The same value also produces a JSON Schema, with the constraints carried over:
Sch.to_json_schema schema
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"name": { "type": "string", "minLength": 2 },
"email": { "type": "string", "format": "email" },
"age": { "type": "integer", "minimum": 17, "maximum": 120 }
},
"required": ["name", "email", "age"]
}
Encoding uses the same schema through Sch.Json.encode_string. This is what Tapak relies on: one definition
validates the request at runtime and describes it in the OpenAPI document.
Sch can do more than this, including optional fields, defaults, cross-field validation, and discriminated unions for variant types. I won’t cover them here. The test suite is currently the most complete reference.
Tradeoffs
Sch currently interprets schemas directly, traversing their structure during encoding, decoding and schema generation. This keeps the implementation straightforward, but repeated interpretation may add overhead.
One possible optimization is to compile a schema into specialized encoder and decoder functions once, then reuse them. I haven’t measured whether this is necessary yet.
For now, the direct interpreter is simple and works well for Tapak’s needs.
Where It Is Now
Sch started as a way to solve Tapak’s validation and OpenAPI generation problem. It eventually became a separate library because the underlying problem isn’t specific to HTTP: the same data structure often needs to be validated, serialized, and described in different formats.
The main lesson for me was that validation and data description don’t have to be separate concerns. By representing constraints as inspectable data, Sch can interpret the same schema in several ways without maintaining separate definitions.
The API is also still evolving as I use it in Tapak and other applications.
The Sch documentation covers installation, supported schemas, and examples. Feedback and contributions are welcome.
Subscribe to the newsletter
Get new articles by email, including posts reserved for members.
We'll send a confirmation link. Open it to subscribe and activate your membership.
Check your inbox
We sent a confirmation link if the address is valid. Open it to subscribe.