mirror of
https://github.com/vinta/awesome-python.git
synced 2026-10-06 17:05:16 +08:00
docs: rewrite Data Validation category intro
The intro predated the audit and never mentioned great-expectations, which the README now lists. Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -1,15 +1,18 @@
|
||||
Declare API input and config as type hints, and Pydantic validates them. Pandera brings Python data validation to dataframes, and jsonschema covers JSON Schema.
|
||||
Pydantic checks API input against your type hints. Python data validation extends to dataframes in Pandera, where each column gets a type hint.
|
||||
|
||||
How to choose:
|
||||
|
||||
- API input, forms, and config: Pydantic
|
||||
- API payloads and other records, typed with hints: Pydantic
|
||||
- Dataframes in pandas, Polars, PySpark, and more: Pandera
|
||||
- Data checked against a JSON Schema document: jsonschema
|
||||
- pandas, polars, or PySpark dataframes: Pandera
|
||||
- Data quality checks on databases and files, with alerts and reports: Great Expectations
|
||||
|
||||
Pydantic builds the schema from your [type hints](https://pydantic.dev/docs/validation/latest/get-started/why/) and guarantees the types of the [output, not the input](https://pydantic.dev/docs/validation/latest/concepts/models/): by default, a numeric string passed to an int field comes out as an int. Where a wrong type should raise an error instead, turn on [strict mode](https://pydantic.dev/docs/validation/latest/concepts/strict_mode/) per field or per model. [Validate incoming JSON directly](https://pydantic.dev/docs/validation/latest/concepts/performance/) instead of parsing it into a dict first. To load config from environment variables, use [Pydantic Settings](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/).
|
||||
Pydantic [takes its schema from Python type hints](https://pydantic.dev/docs/validation/latest/get-started/why/#type-hints), so your type checker and IDE read the same code it validates. Subclass `BaseModel` and annotate its fields, then pass untrusted data to the model: if no `ValidationError` is raised, [every field matches its declared type](https://pydantic.dev/docs/validation/latest/concepts/models/). Your models can also [generate JSON Schema](https://pydantic.dev/docs/validation/latest/get-started/why/#json-schema) for self-documenting APIs and other tools.
|
||||
|
||||
jsonschema is an [implementation of the JSON Schema specification](https://python-jsonschema.readthedocs.io/en/stable/), so one schema can [work across different systems and platforms](https://json-schema.org/overview/what-is-jsonschema). When you validate many instances against one schema, [create a validator for your schema's draft once and call its `validate` method](https://python-jsonschema.readthedocs.io/en/stable/validate/). Use [`iter_errors()`](https://python-jsonschema.readthedocs.io/en/stable/errors/) to report every error, not only the first. To enforce `format` keywords such as dates or emails, [hook a format checker](https://python-jsonschema.readthedocs.io/en/stable/validate/#validating-formats) into the validator.
|
||||
Pandera [validates pandas, Polars, PySpark, and other dataframes against one schema](https://pandera.readthedocs.io/en/latest/). Write that schema as a `DataFrameModel`, [much the way you'd define a Pydantic model](https://pandera.readthedocs.io/en/latest/dataframe_models.html), with a typed field per column. Then decorate your pipeline functions with `check_types()`, which [validates both inputs and outputs](https://pandera.readthedocs.io/en/latest/decorators.html#check-inputs-and-outputs) from their type annotations.
|
||||
|
||||
Pandera validates [dataframe-like objects](https://pandera.readthedocs.io/en/stable/): define a schema once and use it on pandas, polars, PySpark, and other dataframe libraries. Write the schema as a DataFrameModel class, [much like a Pydantic model](https://pandera.readthedocs.io/en/stable/dataframe_models.html), and add the `check_types()` decorator to validate at run time. To check an existing pipeline, put [`check_input()` and `check_output()`](https://pandera.readthedocs.io/en/stable/decorators.html) on its functions. To see every failure in one run instead of only the first, validate with [`lazy=True`](https://pandera.readthedocs.io/en/stable/lazy_validation.html).
|
||||
jsonschema [implements the JSON Schema specification](https://python-jsonschema.readthedocs.io/en/latest/), so the schema is a JSON document rather than Python code. That suits a schema that has to work outside Python too, since JSON Schema [establishes a common language for data exchange](https://json-schema.org/overview/what-is-jsonschema) across systems. Declare `$schema` in each schema to [identify which version it's written for](https://python-jsonschema.readthedocs.io/en/latest/referencing/).
|
||||
|
||||
Pick by the shape of your data. Running a dataframe through a Pydantic model row by row [might not scale](https://pandera.readthedocs.io/en/stable/pydantic_integration.html) to larger datasets, so use Pandera there; a DataFrameModel can still be a field in a Pydantic model. Pydantic can [generate a JSON Schema](https://pydantic.dev/docs/validation/latest/concepts/json_schema/) from any model for tools that read the format. jsonschema works the other way: it validates data against a JSON Schema document you already have.
|
||||
Great Expectations is [a framework for describing data with expressive tests](https://docs.greatexpectations.io/docs/core/introduction/gx_overview/#gx-core-components-and-workflows) and validating data against them, from databases to files in cloud storage. Every script [starts by creating a Data Context](https://docs.greatexpectations.io/docs/core/set_up_a_gx_environment/create_a_data_context/). Group your production Expectations into [Expectation Suites](https://docs.greatexpectations.io/docs/core/define_expectations/organize_expectation_suites/), and run those through [a Checkpoint](https://docs.greatexpectations.io/docs/core/introduction/gx_overview/#run-validations), which can send email or Slack alerts and write the results to Data Docs, a human-readable report.
|
||||
|
||||
Records and tables need different validators, but the picks work together. A Pandera `DataFrameModel` [works as a field on a Pydantic model](https://pandera.readthedocs.io/en/latest/pydantic_integration.html), with Pydantic checking the built-in types next to it. Pandera's maintainer suggests [Pandera for in-memory dataframes and Great Expectations for data on disk](https://github.com/unionai-oss/pandera/discussions/598). Pandera needs zero configuration, while Great Expectations takes some setup up front.
|
||||
|
||||
Reference in New Issue
Block a user