JSON Schema
Declaring what your JSON must contain — so that malformed data is caught at the boundary instead of three layers into your application.
The Gap Schema Fills
A syntax validator answers one question: can this text be parsed? That is a low bar. This document clears it and would still break almost any application that received it:
{
"userId": "not-a-number",
"email": null,
"age": -5,
"status": "bananas",
"roles": "admin"
}Every bracket matches and every string is quoted, so a parser accepts it happily. But userId is a string where a number was expected, age is negative, status holds a value outside the permitted set, and roles is a bare string where an array belongs. Without a schema, each of these surfaces later — as a type error deep in your code, or worse, as silently wrong data written to your database.
JSON Schema lets you state the expectation once, machine-readably, and reject bad input at the boundary.
A Schema Is Just JSON
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "User",
"type": "object",
"required": ["userId", "email"],
"properties": {
"userId": { "type": "integer", "minimum": 1 },
"email": { "type": "string", "format": "email" },
"age": { "type": "integer", "minimum": 0, "maximum": 150 },
"status": { "enum": ["active", "suspended", "deleted"] },
"roles": {
"type": "array",
"items": { "type": "string" },
"minItems": 1,
"uniqueItems": true
}
},
"additionalProperties": false
}Reading it top to bottom: the document must be an object; userId and email must be present; userId must be a positive integer; status must be one of three values; roles must be a non-empty array of distinct strings; and no properties beyond those listed are permitted.
Because the schema is itself JSON, it travels anywhere JSON travels — checked into your repository, served from an endpoint, shared with a team writing in a different language, or embedded in an OpenAPI document.
The Keywords Worth Knowing
| Keyword | Applies to | Meaning |
|---|---|---|
| type | any | One of the JSON types, or a list of them |
| required | object | Keys that must be present |
| enum / const | any | Restrict to a fixed set, or one exact value |
| minimum / maximum | number | Inclusive numeric bounds |
| minLength / pattern | string | Length floor, or a regular expression |
| format | string | Named formats: email, uri, date-time, uuid |
| items | array | Schema every element must satisfy |
| minItems / uniqueItems | array | Length floor, and no duplicates |
| additionalProperties | object | Whether unlisted keys are allowed |
One caveat on format: in most validators it is annotative by default and does not enforce anything unless you turn enforcement on. In Ajv, for instance, you must install and register ajv-formats. A schema that looks like it validates email addresses and silently does not is a common and unpleasant surprise.
Nesting and Reuse
Real schemas describe nested structures, and $defs with $ref keeps them from repeating themselves:
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"required": ["data"],
"properties": {
"data": {
"type": "array",
"items": { "$ref": "#/$defs/user" }
},
"meta": { "$ref": "#/$defs/pagination" }
},
"$defs": {
"user": {
"type": "object",
"required": ["id", "email"],
"properties": {
"id": { "type": "integer" },
"email": { "type": "string", "format": "email" },
"address": { "$ref": "#/$defs/address" }
}
},
"address": {
"type": "object",
"properties": {
"city": { "type": "string" },
"country": { "type": "string", "pattern": "^[A-Z]{2}quot; }
}
},
"pagination": {
"type": "object",
"required": ["page", "total"],
"properties": {
"page": { "type": "integer", "minimum": 1 },
"total": { "type": "integer", "minimum": 0 }
}
}
}
}$ref can also point at another file or a URL, which is how organisations share a common definition of "user" or "money" across many services.
Nullable Fields and Optionality
A frequent source of confusion. required controls whether a key must be present; it says nothing about whether the value may be null. These are independent axes.
// Present, but may be null
"middleName": { "type": ["string", "null"] }
// May be absent entirely — simply omit it from "required"
// Must be present AND must not be null
"email": { "type": "string" } // with "email" listed in requiredDeciding this deliberately matters more than which option you pick. As covered on the data types page, null and "absent" carry different meanings — particularly for PATCH requests, where null means "clear this" and omission means "leave it alone".
additionalProperties: Two Different Answers
Setting additionalProperties: false rejects any key the schema does not list. Whether that is correct depends entirely on which direction the data is travelling.
Inbound requests — usually strict
A client sending emial instead of email has made a mistake. Silently ignoring it stores an empty address and reports success. Rejecting it surfaces the bug immediately.
Consumed responses — usually lenient
If you validate a third party's responses strictly, their adding a harmless new field breaks your client on their next deploy. Ignore what you do not recognise.
This is the same trade-off as Jackson's FAIL_ON_UNKNOWN_PROPERTIES in Java and Go's DisallowUnknownFields(). Strict on the way in, tolerant on the way out.
Validating in Practice
// JavaScript — Ajv
import Ajv from "ajv";
import addFormats from "ajv-formats";
const ajv = new Ajv({ allErrors: true });
addFormats(ajv); // required for "format" to enforce
const validate = ajv.compile(schema);
if (!validate(data)) {
console.error(validate.errors); // path, keyword and message per failure
}# Python — jsonschema
from jsonschema import validate, ValidationError
try:
validate(instance=data, schema=schema)
except ValidationError as e:
print(f"{'.'.join(str(p) for p in e.absolute_path)}: {e.message}")Compile the schema once and reuse the validator. Ajv in particular generates optimised validation code at compile time, so recompiling per request throws away most of its performance advantage.
Where Schemas Earn Their Cost
- API boundaries. Reject malformed requests before they touch business logic, with a field-level error message the client can act on.
- Configuration files. Point
$schemaat a URL and editors like VS Code give you autocomplete and inline errors for free. - Contract testing. Assert in CI that your responses still match the published schema, catching accidental breaking changes before release.
- Documentation that cannot drift. A schema is executable, so unlike prose it fails loudly when the implementation changes.
- Code generation. Many toolchains generate types or classes from a schema, keeping several languages consistent from one source.
Frequently Asked Questions
What is JSON Schema?
JSON Schema is a vocabulary, itself written in JSON, for describing the structure a JSON document must have. It lets you declare required fields, value types, permitted ranges and nested shapes, and then check any document against that declaration automatically.
Is JSON Schema part of the JSON specification?
No. JSON itself defines only syntax. JSON Schema is a separate standard maintained by its own working group, currently at draft 2020-12. Support is provided by libraries rather than by language standard libraries.
What is the difference between JSON Schema and OpenAPI?
OpenAPI describes an entire HTTP API — endpoints, methods, authentication, responses — and uses JSON Schema to describe the shape of the request and response bodies within it. JSON Schema is the data-shape layer; OpenAPI is the whole API contract around it.
Does additionalProperties: false break forward compatibility?
It can. Rejecting unknown fields is right for inbound requests, where an unrecognised property usually means a client typo. It is usually wrong for validating responses you consume, because the provider adding a harmless new field would break your client on their next deploy.
Which library should I use?
Ajv for JavaScript and TypeScript, jsonschema or fastjsonschema for Python, networknt json-schema-validator for Java, and santhosh-tekuri/jsonschema for Go. If you only need runtime validation in TypeScript and do not have to share the schema across languages, Zod is often more ergonomic.