diff --git a/website/data/category_intros/data-analysis.md b/website/data/category_intros/data-analysis.md new file mode 100644 index 00000000..5d640802 --- /dev/null +++ b/website/data/category_intros/data-analysis.md @@ -0,0 +1,17 @@ +Once your data nears the size of your RAM, switch Python data analysis libraries from pandas to Polars. Once it lives in a database, Ibis runs your code there. + +How to choose: + +- Data that fits in memory: pandas +- Handing DataFrames to libraries that expect pandas: pandas +- Large data on one machine, even bigger than your RAM: Polars +- Data already in a database or warehouse: Ibis +- Code you prototype locally, then run on a warehouse or cluster: Ibis + +pandas aims to be [the fundamental high-level building block](https://pandas.pydata.org/docs/getting_started/overview.html) for practical data analysis in Python. Its DataFrame fits time series and tables with mixed column types, like an SQL table or a spreadsheet. It's built on NumPy to work with the rest of the scientific Python stack. In production code, the docs [recommend `.loc` and `.iloc` over `[]`](https://pandas.pydata.org/docs/user_guide/indexing.html) to select data. Set values in [one `.loc` call](https://pandas.pydata.org/docs/user_guide/copy_on_write.html#chained-assignment), not chained indexing like `df["foo"][mask] = 100`. pandas keeps everything in memory, so a dataset that takes a sizable share of your RAM [gets unwieldy](https://pandas.pydata.org/docs/user_guide/scale.html), and the docs point you to other libraries. + +Polars is [built for multithreaded computing on a single machine](https://docs.pola.rs/user-guide/misc/comparison/#pandas), and its stricter API leads to fewer schema bugs. It has no index: rows are known by their position, so [no index state can change what a query means](https://docs.pola.rs/user-guide/migration/pandas/#polars-does-not-have-a-multi-indexindex). Make the lazy API [your default](https://docs.pola.rs/user-guide/migration/pandas/#be-lazy): start from a function like `scan_csv()` or call `.lazy()`, then `.collect()` the result. That lets the optimizer [filter rows and pick columns while reading the data](https://docs.pola.rs/user-guide/concepts/lazy-api/). Stay eager for exploratory work, when [you don't know yet what your query will look like](https://docs.pola.rs/user-guide/concepts/lazy-api/#when-to-use-which). Write expressions inside `select`, `with_columns`, `filter`, and `group_by`, since Polars code that [looks like pandas code](https://docs.pola.rs/user-guide/migration/pandas/#key-syntax-differences) likely runs slower than it should. For data that doesn't fit in memory, run the query on the [streaming engine](https://docs.pola.rs/user-guide/concepts/streaming/), which works through it in batches. + +Ibis [compiles one Python dataframe API](https://ibis-project.org/why#how-does-ibis-work) into each backend's native language, mostly SQL, so the database or engine does the work. pandas' ecosystem page lists it for [bridging local Python and remote databases](https://pandas.pydata.org/community/ecosystem.html#ibis). Start on the default DuckDB backend, as the tutorial [recommends](https://ibis-project.org/tutorials/basics#install-ibis), then [change the connection string](https://ibis-project.org/why#scaling-up-and-out) to run the same code on PySpark, BigQuery, or Trino. Expressions are lazy: nothing runs until you call a method like `to_pandas()`, and only then does Ibis [send the compiled query](https://ibis-project.org/tutorials/coming-from/pandas) to the backend. + +You can mix all three. Polars converts a DataFrame [with `to_pandas()`](https://docs.pola.rs/api/python/stable/reference/dataframe/api/polars.DataFrame.to_pandas.html), Ibis returns results [as pandas or Polars DataFrames](https://ibis-project.org/reference/expression-tables#ibis.expr.types.relations.Table.to_pandas), and pandas [can use PyArrow](https://pandas.pydata.org/docs/user_guide/pyarrow.html) to trade data with Arrow-based libraries like Polars. Do the heavy work in Polars or Ibis, and hand a pandas DataFrame to the libraries that expect one.