mirror of
https://github.com/vinta/awesome-python.git
synced 2026-10-06 17:05:16 +08:00
docs: rewrite Data Validation category intro
Covers Pydantic for API input and config, Pandera for dataframes, and jsonschema for JSON Schema validation. Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -1,17 +1,15 @@
|
||||
For a Python validation library, use Pydantic for data coming into your app, jsonschema when a JSON Schema is the contract, and Pandera for dataframes.
|
||||
Validate API input and config with Pydantic, the Python data validation library built on type hints. Use Pandera for dataframes, jsonschema for JSON Schema.
|
||||
|
||||
How to choose:
|
||||
|
||||
- API payloads, config files, anything you can describe with type hints: Pydantic
|
||||
- A JSON Schema shared with other languages or teams: jsonschema
|
||||
- pandas, Polars, or PySpark dataframes: Pandera
|
||||
- API input, forms, and config: Pydantic
|
||||
- Data checked against a JSON Schema document: jsonschema
|
||||
- pandas, polars, or PySpark dataframes: Pandera
|
||||
|
||||
With Pydantic, describe your data as a `BaseModel` with type hints. Parse JSON with `model_validate_json()`, not `model_validate(json.loads(...))`, as [its performance tips](https://pydantic.dev/docs/validation/latest/concepts/performance/) recommend. For a type that isn't a model, like `list[Item]`, create one `TypeAdapter` and reuse it.
|
||||
Pydantic builds the schema from your [type hints](https://pydantic.dev/docs/validation/latest/get-started/why/) and guarantees the types of the [output, not the input](https://pydantic.dev/docs/validation/latest/concepts/models/): by default, a numeric string passed to an int field comes out as an int. Where a wrong type should raise an error instead, turn on [strict mode](https://pydantic.dev/docs/validation/latest/concepts/strict_mode/) per field or per model. [Validate incoming JSON directly](https://pydantic.dev/docs/validation/latest/concepts/performance/) instead of parsing it into a dict first. To load config from environment variables, use [Pydantic Settings](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/).
|
||||
|
||||
Pydantic is forgiving by default: it turns `"123"` into `123` and ignores fields your model doesn't declare. When that's too loose, turn on [strict mode](https://pydantic.dev/docs/validation/latest/concepts/strict_mode/) or set `extra='forbid'`.
|
||||
jsonschema is an [implementation of the JSON Schema specification](https://python-jsonschema.readthedocs.io/en/stable/), so one schema can [work across different systems and platforms](https://json-schema.org/overview/what-is-jsonschema). When you validate many instances against one schema, [create a validator for your schema's draft once and call its `validate` method](https://python-jsonschema.readthedocs.io/en/stable/validate/). Use [`iter_errors()`](https://python-jsonschema.readthedocs.io/en/stable/errors/) to report every error, not only the first. To enforce `format` keywords such as dates or emails, [hook a format checker](https://python-jsonschema.readthedocs.io/en/stable/validate/#validating-formats) into the validator.
|
||||
|
||||
With jsonschema, `validate()` checks the schema itself on every call. To validate many documents against one schema, [create a validator once](https://python-jsonschema.readthedocs.io/en/stable/validate/) and reuse it, like `Draft202012Validator(schema)`. The `format` keyword checks nothing until you pass a `format_checker`, and [`default` doesn't fill in missing fields](https://python-jsonschema.readthedocs.io/en/stable/faq/).
|
||||
Pandera validates [dataframe-like objects](https://pandera.readthedocs.io/en/stable/): define a schema once and use it on pandas, polars, PySpark, and other dataframe libraries. Write the schema as a DataFrameModel class, [much like a Pydantic model](https://pandera.readthedocs.io/en/stable/dataframe_models.html), and add the `check_types()` decorator to validate at run time. To check an existing pipeline, put [`check_input()` and `check_output()`](https://pandera.readthedocs.io/en/stable/decorators.html) on its functions. To see every failure in one run instead of only the first, validate with [`lazy=True`](https://pandera.readthedocs.io/en/stable/lazy_validation.html).
|
||||
|
||||
With Pandera, write a `DataFrameModel`, which [works like a Pydantic model](https://pandera.readthedocs.io/en/stable/dataframe_models.html) for your columns. Decorate pipeline functions with `@pa.check_types` to validate what goes in and what comes out. Pass `lazy=True` to [get every failure in one report](https://pandera.readthedocs.io/en/stable/lazy_validation.html), not just the first.
|
||||
|
||||
Validate once, where untrusted data enters your program: a request body, a config file, a CSV upload. After that, [Pydantic guarantees](https://pydantic.dev/docs/validation/latest/concepts/models/) the fields match their types, so don't check them again deeper in. The same goes for schemas: if other teams need a JSON Schema, [generate it from your Pydantic models](https://pydantic.dev/docs/validation/latest/concepts/json_schema/) instead of writing it twice.
|
||||
Pick by the shape of your data. Running a dataframe through a Pydantic model row by row [might not scale](https://pandera.readthedocs.io/en/stable/pydantic_integration.html) to larger datasets, so use Pandera there; a DataFrameModel can still be a field in a Pydantic model. Pydantic can [generate a JSON Schema](https://pydantic.dev/docs/validation/latest/concepts/json_schema/) from any model for tools that read the format. jsonschema works the other way: it validates data against a JSON Schema document you already have.
|
||||
|
||||
Reference in New Issue
Block a user