docs: rewrite Data Validation category intro

Covers Pydantic for API input and config, Pandera for dataframes, and jsonschema for JSON Schema validation.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Vinta Chen
2026-09-27 00:27:34 +08:00
co-authored by Claude
parent c52f21f99d
commit e5088a6f5c
@@ -1,17 +1,15 @@
For a Python validation library, use Pydantic for data coming into your app, jsonschema when a JSON Schema is the contract, and Pandera for dataframes.
Validate API input and config with Pydantic, the Python data validation library built on type hints. Use Pandera for dataframes, jsonschema for JSON Schema.
How to choose:
- API payloads, config files, anything you can describe with type hints: Pydantic
- A JSON Schema shared with other languages or teams: jsonschema
- pandas, Polars, or PySpark dataframes: Pandera
- API input, forms, and config: Pydantic
- Data checked against a JSON Schema document: jsonschema
- pandas, polars, or PySpark dataframes: Pandera
With Pydantic, describe your data as a `BaseModel` with type hints. Parse JSON with `model_validate_json()`, not `model_validate(json.loads(...))`, as [its performance tips](https://pydantic.dev/docs/validation/latest/concepts/performance/) recommend. For a type that isn't a model, like `list[Item]`, create one `TypeAdapter` and reuse it.
Pydantic builds the schema from your [type hints](https://pydantic.dev/docs/validation/latest/get-started/why/) and guarantees the types of the [output, not the input](https://pydantic.dev/docs/validation/latest/concepts/models/): by default, a numeric string passed to an int field comes out as an int. Where a wrong type should raise an error instead, turn on [strict mode](https://pydantic.dev/docs/validation/latest/concepts/strict_mode/) per field or per model. [Validate incoming JSON directly](https://pydantic.dev/docs/validation/latest/concepts/performance/) instead of parsing it into a dict first. To load config from environment variables, use [Pydantic Settings](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/).
Pydantic is forgiving by default: it turns `"123"` into `123` and ignores fields your model doesn't declare. When that's too loose, turn on [strict mode](https://pydantic.dev/docs/validation/latest/concepts/strict_mode/) or set `extra='forbid'`.
jsonschema is an [implementation of the JSON Schema specification](https://python-jsonschema.readthedocs.io/en/stable/), so one schema can [work across different systems and platforms](https://json-schema.org/overview/what-is-jsonschema). When you validate many instances against one schema, [create a validator for your schema's draft once and call its `validate` method](https://python-jsonschema.readthedocs.io/en/stable/validate/). Use [`iter_errors()`](https://python-jsonschema.readthedocs.io/en/stable/errors/) to report every error, not only the first. To enforce `format` keywords such as dates or emails, [hook a format checker](https://python-jsonschema.readthedocs.io/en/stable/validate/#validating-formats) into the validator.
With jsonschema, `validate()` checks the schema itself on every call. To validate many documents against one schema, [create a validator once](https://python-jsonschema.readthedocs.io/en/stable/validate/) and reuse it, like `Draft202012Validator(schema)`. The `format` keyword checks nothing until you pass a `format_checker`, and [`default` doesn't fill in missing fields](https://python-jsonschema.readthedocs.io/en/stable/faq/).
Pandera validates [dataframe-like objects](https://pandera.readthedocs.io/en/stable/): define a schema once and use it on pandas, polars, PySpark, and other dataframe libraries. Write the schema as a DataFrameModel class, [much like a Pydantic model](https://pandera.readthedocs.io/en/stable/dataframe_models.html), and add the `check_types()` decorator to validate at run time. To check an existing pipeline, put [`check_input()` and `check_output()`](https://pandera.readthedocs.io/en/stable/decorators.html) on its functions. To see every failure in one run instead of only the first, validate with [`lazy=True`](https://pandera.readthedocs.io/en/stable/lazy_validation.html).
With Pandera, write a `DataFrameModel`, which [works like a Pydantic model](https://pandera.readthedocs.io/en/stable/dataframe_models.html) for your columns. Decorate pipeline functions with `@pa.check_types` to validate what goes in and what comes out. Pass `lazy=True` to [get every failure in one report](https://pandera.readthedocs.io/en/stable/lazy_validation.html), not just the first.
Validate once, where untrusted data enters your program: a request body, a config file, a CSV upload. After that, [Pydantic guarantees](https://pydantic.dev/docs/validation/latest/concepts/models/) the fields match their types, so don't check them again deeper in. The same goes for schemas: if other teams need a JSON Schema, [generate it from your Pydantic models](https://pydantic.dev/docs/validation/latest/concepts/json_schema/) instead of writing it twice.
Pick by the shape of your data. Running a dataframe through a Pydantic model row by row [might not scale](https://pandera.readthedocs.io/en/stable/pydantic_integration.html) to larger datasets, so use Pandera there; a DataFrameModel can still be a field in a Pydantic model. Pydantic can [generate a JSON Schema](https://pydantic.dev/docs/validation/latest/concepts/json_schema/) from any model for tools that read the format. jsonschema works the other way: it validates data against a JSON Schema document you already have.