mirror of
https://github.com/vinta/awesome-python.git
synced 2026-10-02 08:23:10 +08:00
docs: add Data Analysis category intro
Data Analysis's category page had no intro, so its meta description fell back to generic text and gave readers no guidance on which library to pick. Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,17 @@
|
||||
Once your data nears the size of your RAM, switch Python data analysis libraries from pandas to Polars. Once it lives in a database, Ibis runs your code there.
|
||||
|
||||
How to choose:
|
||||
|
||||
- Data that fits in memory: pandas
|
||||
- Handing DataFrames to libraries that expect pandas: pandas
|
||||
- Large data on one machine, even bigger than your RAM: Polars
|
||||
- Data already in a database or warehouse: Ibis
|
||||
- Code you prototype locally, then run on a warehouse or cluster: Ibis
|
||||
|
||||
pandas aims to be [the fundamental high-level building block](https://pandas.pydata.org/docs/getting_started/overview.html) for practical data analysis in Python. Its DataFrame fits time series and tables with mixed column types, like an SQL table or a spreadsheet. It's built on NumPy to work with the rest of the scientific Python stack. In production code, the docs [recommend `.loc` and `.iloc` over `[]`](https://pandas.pydata.org/docs/user_guide/indexing.html) to select data. Set values in [one `.loc` call](https://pandas.pydata.org/docs/user_guide/copy_on_write.html#chained-assignment), not chained indexing like `df["foo"][mask] = 100`. pandas keeps everything in memory, so a dataset that takes a sizable share of your RAM [gets unwieldy](https://pandas.pydata.org/docs/user_guide/scale.html), and the docs point you to other libraries.
|
||||
|
||||
Polars is [built for multithreaded computing on a single machine](https://docs.pola.rs/user-guide/misc/comparison/#pandas), and its stricter API leads to fewer schema bugs. It has no index: rows are known by their position, so [no index state can change what a query means](https://docs.pola.rs/user-guide/migration/pandas/#polars-does-not-have-a-multi-indexindex). Make the lazy API [your default](https://docs.pola.rs/user-guide/migration/pandas/#be-lazy): start from a function like `scan_csv()` or call `.lazy()`, then `.collect()` the result. That lets the optimizer [filter rows and pick columns while reading the data](https://docs.pola.rs/user-guide/concepts/lazy-api/). Stay eager for exploratory work, when [you don't know yet what your query will look like](https://docs.pola.rs/user-guide/concepts/lazy-api/#when-to-use-which). Write expressions inside `select`, `with_columns`, `filter`, and `group_by`, since Polars code that [looks like pandas code](https://docs.pola.rs/user-guide/migration/pandas/#key-syntax-differences) likely runs slower than it should. For data that doesn't fit in memory, run the query on the [streaming engine](https://docs.pola.rs/user-guide/concepts/streaming/), which works through it in batches.
|
||||
|
||||
Ibis [compiles one Python dataframe API](https://ibis-project.org/why#how-does-ibis-work) into each backend's native language, mostly SQL, so the database or engine does the work. pandas' ecosystem page lists it for [bridging local Python and remote databases](https://pandas.pydata.org/community/ecosystem.html#ibis). Start on the default DuckDB backend, as the tutorial [recommends](https://ibis-project.org/tutorials/basics#install-ibis), then [change the connection string](https://ibis-project.org/why#scaling-up-and-out) to run the same code on PySpark, BigQuery, or Trino. Expressions are lazy: nothing runs until you call a method like `to_pandas()`, and only then does Ibis [send the compiled query](https://ibis-project.org/tutorials/coming-from/pandas) to the backend.
|
||||
|
||||
You can mix all three. Polars converts a DataFrame [with `to_pandas()`](https://docs.pola.rs/api/python/stable/reference/dataframe/api/polars.DataFrame.to_pandas.html), Ibis returns results [as pandas or Polars DataFrames](https://ibis-project.org/reference/expression-tables#ibis.expr.types.relations.Table.to_pandas), and pandas [can use PyArrow](https://pandas.pydata.org/docs/user_guide/pyarrow.html) to trade data with Arrow-based libraries like Polars. Do the heavy work in Polars or Ibis, and hand a pandas DataFrame to the libraries that expect one.
|
||||
Reference in New Issue
Block a user