docs: rewrite Data Validation category intro

The intro predated the audit and never mentioned great-expectations, which the README now lists.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Vinta Chen
2026-10-02 16:37:46 +08:00
co-authored by Claude
parent 9a6bfac9d9
commit 40cacb9319
@@ -1,15 +1,18 @@
Declare API input and config as type hints, and Pydantic validates them. Pandera brings Python data validation to dataframes, and jsonschema covers JSON Schema.
Pydantic checks API input against your type hints. Python data validation extends to dataframes in Pandera, where each column gets a type hint.
How to choose:
- API input, forms, and config: Pydantic
- API payloads and other records, typed with hints: Pydantic
- Dataframes in pandas, Polars, PySpark, and more: Pandera
- Data checked against a JSON Schema document: jsonschema
- pandas, polars, or PySpark dataframes: Pandera
- Data quality checks on databases and files, with alerts and reports: Great Expectations
Pydantic builds the schema from your [type hints](https://pydantic.dev/docs/validation/latest/get-started/why/) and guarantees the types of the [output, not the input](https://pydantic.dev/docs/validation/latest/concepts/models/): by default, a numeric string passed to an int field comes out as an int. Where a wrong type should raise an error instead, turn on [strict mode](https://pydantic.dev/docs/validation/latest/concepts/strict_mode/) per field or per model. [Validate incoming JSON directly](https://pydantic.dev/docs/validation/latest/concepts/performance/) instead of parsing it into a dict first. To load config from environment variables, use [Pydantic Settings](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/).
Pydantic [takes its schema from Python type hints](https://pydantic.dev/docs/validation/latest/get-started/why/#type-hints), so your type checker and IDE read the same code it validates. Subclass `BaseModel` and annotate its fields, then pass untrusted data to the model: if no `ValidationError` is raised, [every field matches its declared type](https://pydantic.dev/docs/validation/latest/concepts/models/). Your models can also [generate JSON Schema](https://pydantic.dev/docs/validation/latest/get-started/why/#json-schema) for self-documenting APIs and other tools.
jsonschema is an [implementation of the JSON Schema specification](https://python-jsonschema.readthedocs.io/en/stable/), so one schema can [work across different systems and platforms](https://json-schema.org/overview/what-is-jsonschema). When you validate many instances against one schema, [create a validator for your schema's draft once and call its `validate` method](https://python-jsonschema.readthedocs.io/en/stable/validate/). Use [`iter_errors()`](https://python-jsonschema.readthedocs.io/en/stable/errors/) to report every error, not only the first. To enforce `format` keywords such as dates or emails, [hook a format checker](https://python-jsonschema.readthedocs.io/en/stable/validate/#validating-formats) into the validator.
Pandera [validates pandas, Polars, PySpark, and other dataframes against one schema](https://pandera.readthedocs.io/en/latest/). Write that schema as a `DataFrameModel`, [much the way you'd define a Pydantic model](https://pandera.readthedocs.io/en/latest/dataframe_models.html), with a typed field per column. Then decorate your pipeline functions with `check_types()`, which [validates both inputs and outputs](https://pandera.readthedocs.io/en/latest/decorators.html#check-inputs-and-outputs) from their type annotations.
Pandera validates [dataframe-like objects](https://pandera.readthedocs.io/en/stable/): define a schema once and use it on pandas, polars, PySpark, and other dataframe libraries. Write the schema as a DataFrameModel class, [much like a Pydantic model](https://pandera.readthedocs.io/en/stable/dataframe_models.html), and add the `check_types()` decorator to validate at run time. To check an existing pipeline, put [`check_input()` and `check_output()`](https://pandera.readthedocs.io/en/stable/decorators.html) on its functions. To see every failure in one run instead of only the first, validate with [`lazy=True`](https://pandera.readthedocs.io/en/stable/lazy_validation.html).
jsonschema [implements the JSON Schema specification](https://python-jsonschema.readthedocs.io/en/latest/), so the schema is a JSON document rather than Python code. That suits a schema that has to work outside Python too, since JSON Schema [establishes a common language for data exchange](https://json-schema.org/overview/what-is-jsonschema) across systems. Declare `$schema` in each schema to [identify which version it's written for](https://python-jsonschema.readthedocs.io/en/latest/referencing/).
Pick by the shape of your data. Running a dataframe through a Pydantic model row by row [might not scale](https://pandera.readthedocs.io/en/stable/pydantic_integration.html) to larger datasets, so use Pandera there; a DataFrameModel can still be a field in a Pydantic model. Pydantic can [generate a JSON Schema](https://pydantic.dev/docs/validation/latest/concepts/json_schema/) from any model for tools that read the format. jsonschema works the other way: it validates data against a JSON Schema document you already have.
Great Expectations is [a framework for describing data with expressive tests](https://docs.greatexpectations.io/docs/core/introduction/gx_overview/#gx-core-components-and-workflows) and validating data against them, from databases to files in cloud storage. Every script [starts by creating a Data Context](https://docs.greatexpectations.io/docs/core/set_up_a_gx_environment/create_a_data_context/). Group your production Expectations into [Expectation Suites](https://docs.greatexpectations.io/docs/core/define_expectations/organize_expectation_suites/), and run those through [a Checkpoint](https://docs.greatexpectations.io/docs/core/introduction/gx_overview/#run-validations), which can send email or Slack alerts and write the results to Data Docs, a human-readable report.
Records and tables need different validators, but the picks work together. A Pandera `DataFrameModel` [works as a field on a Pydantic model](https://pandera.readthedocs.io/en/latest/pydantic_integration.html), with Pydantic checking the built-in types next to it. Pandera's maintainer suggests [Pandera for in-memory dataframes and Great Expectations for data on disk](https://github.com/unionai-oss/pandera/discussions/598). Pandera needs zero configuration, while Great Expectations takes some setup up front.