diff --git a/website/data/category_intros/science.md b/website/data/category_intros/science.md new file mode 100644 index 00000000..87f70c05 --- /dev/null +++ b/website/data/category_intros/science.md @@ -0,0 +1,42 @@ +Underneath most Python scientific computing libraries sits NumPy's array. SciPy adds algorithms on top, and Numba compiles loops that won't vectorize. + +How to choose: + +- Arrays and vectorized math: NumPy +- Optimization, integration, interpolation, and linear algebra: SciPy +- Numeric loops that won't vectorize: Numba +- Exact, symbolic math: SymPy +- Regression, time series, and hypothesis tests: statsmodels +- Biological sequences and file formats: Biopython; molecules: RDKit +- Physical units: Pint; astronomy: Astropy; seismology: ObsPy +- Bayesian models: PyMC +- Discrete-event simulations: SimPy; agent-based models: Mesa +- Graphs and networks: NetworkX +- Planar geometry: Shapely +- Colour science: Colour; math animations: Manim + +NumPy gives you a multidimensional array and fast routines that run on it. Write whole-array expressions instead of Python loops: [vectorization and broadcasting](https://numpy.org/doc/stable/user/whatisnumpy.html#why-is-numpy-fast) push the element-by-element work into compiled C code, and your code reads closer to the math. For random numbers, [create a Generator with `default_rng()`](https://numpy.org/doc/stable/reference/random/index.html#random-quick-start) and call its methods. + +SciPy is [a collection of algorithms built on NumPy](https://docs.scipy.org/doc/scipy/tutorial/index.html), split into subpackages by domain, like `optimize`, `integrate`, `interpolate`, `linalg`, and `stats`. Its docs recommend [using those subpackages as namespaces](https://docs.scipy.org/doc/scipy/reference/index.html#guidelines-for-importing-functions-from-scipy): `from scipy import optimize`, then `optimize.curve_fit()`. For linear algebra, use `scipy.linalg` over `numpy.linalg`. It has everything NumPy's module has and more, and it's [always compiled with BLAS/LAPACK](https://docs.scipy.org/doc/scipy/tutorial/linalg.html#scipy-linalg-vs-numpy-linalg), so it may run faster. + +Numba is a just-in-time compiler that [works best on code that uses NumPy arrays and functions, and loops](https://numba.readthedocs.io/en/stable/user/5minguide.html). NumPy wants vector operations, but Numba is [happy with plain loops](https://numba.readthedocs.io/en/stable/user/performance-tips.html#loops), so it fits the numeric code you can't vectorize. Put `@jit` on the function and [let Numba decide when and how to optimize](https://numba.readthedocs.io/en/stable/user/jit.html#lazy-compilation). Keep pandas out of those functions: Numba doesn't understand it, so that code [runs in the interpreter](https://numba.readthedocs.io/en/stable/user/5minguide.html#will-numba-work-for-my-code) with Numba's overhead on top. + +SymPy does symbolic math: expressions stay [exact, not approximate](https://docs.sympy.org/latest/tutorials/intro-tutorial/intro.html#what-is-symbolic-computation), and it's written entirely in Python. Its best practices keep symbolic and numeric code apart: [model the problem in SymPy, then turn the result into a function with `lambdify()`](https://docs.sympy.org/latest/explanation/best-practices.html#separate-symbolic-and-numeric-code) that runs on NumPy arrays. Define symbols with `symbols()` and [the assumptions you know](https://docs.sympy.org/latest/explanation/best-practices.html#defining-symbols), like `positive=True`, so more expressions simplify. Build expressions from those symbols, [not from strings](https://docs.sympy.org/latest/explanation/best-practices.html#avoid-string-inputs). + +statsmodels [estimates statistical models, runs hypothesis tests, and explores data](https://www.statsmodels.org/stable/index.html), and [most of its results are verified](https://www.statsmodels.org/stable/about.html#testing) against another statistical package. SciPy's stats docs [send regression, linear models, and time series analysis to statsmodels](https://docs.scipy.org/doc/scipy/reference/stats.html), while `scipy.stats` keeps the probability distributions and statistical tests. Write your model as an [R-style formula](https://www.statsmodels.org/stable/example_formulas.html) on a pandas DataFrame, like `smf.ols("y ~ x", data=df).fit()`, and read the fit with `summary()`. + +Biopython [parses bioinformatics file formats](https://biopython.org/docs/latest/Tutorial/chapter_introduction.html) like FASTA and GenBank, queries online services like NCBI, and gives you a standard sequence class. Read sequence files with `Bio.SeqIO.parse()`, or [`Bio.SeqIO.read()` when a file holds one record](https://biopython.org/docs/latest/Tutorial/chapter_seqio.html#parsing-or-reading-sequences). When you query NCBI through `Bio.Entrez`, [pass your email](https://biopython.org/docs/latest/Tutorial/chapter_entrez.html#entrez-guidelines) so NCBI can contact you if there's a problem. RDKit is a [cheminformatics toolkit with its core in C++](https://www.rdkit.org/docs/Overview.html), for 2D and 3D molecular operations and descriptors for machine learning. To compare molecules, [create a fingerprint generator](https://www.rdkit.org/docs/GettingStartedInPython.html#fingerprinting-and-molecular-similarity) for the fingerprint type you want, the docs' most consistent way to get fingerprints. + +Pint handles [physical quantities](https://pint.readthedocs.io/en/stable/getting/overview.html): a value times a unit, in any numeric type. In a package, [create one `UnitRegistry` in a single place](https://pint.readthedocs.io/en/stable/getting/pint-in-your-projects.html#having-a-shared-registry) and import it everywhere, since quantities from different registries don't mix. Astropy is [a common core package for astronomy](https://www.astropy.org/), with affiliated packages around it. [Import the subpackage you need](https://docs.astropy.org/en/stable/importing_astropy.html#importing-astropy-and-sub-packages), like `from astropy import units as u`, and never import with `*`. ObsPy is [a framework for processing seismological data](https://docs.obspy.org/): `read()` loads [SAC, MiniSEED, and other formats into a Stream](https://docs.obspy.org/tutorial/code_snippets/reading_seismograms.html) of Traces, and its [FDSN client](https://docs.obspy.org/packages/obspy.clients.fdsn.html#basic-fdsn-client-usage) fetches data from data centers. + +PyMC [builds Bayesian models with a simple Python API](https://www.pymc.io/welcome.html) and fits them with Markov chain Monte Carlo or variational inference. Declare the variables inside a `with pm.Model():` block, which [adds them to the model for you](https://www.pymc.io/projects/docs/en/stable/learn/core_notebooks/pymc_overview.html#model-specification), and let `pm.sample()` pick the samplers. Validate the model with [posterior predictive checks](https://www.pymc.io/projects/docs/en/stable/learn/core_notebooks/posterior_predictive.html), and run prior predictive checks too, which its docs call a crucial part of the Bayesian workflow. + +SimPy runs discrete-event simulations. Each process is a Python generator that [yields events and waits for them](https://simpy.readthedocs.io/en/latest/simpy_intro/basic_concepts.html), and shared resources model congestion points like servers and checkout counters. Request a resource in a `with` block so it's [released for you](https://simpy.readthedocs.io/en/latest/simpy_intro/shared_resources.html#basic-resource-usage). Its docs call it [overkill for fixed-step simulations](https://simpy.readthedocs.io/en/latest/) whose processes don't interact or share resources. Mesa builds [agent-based models](https://mesa.readthedocs.io/latest/) from core components like spatial grids, shows them in the browser, and hands the results to Python's data tools. Its best practices put [the model class in `model.py` and the agents in `agents.py`](https://mesa.readthedocs.io/latest/best-practices.html#model-layout), with an optional browser visualization in `app.py`. Its [DataCollector](https://mesa.readthedocs.io/latest/getting_started.html) collects model-level and agent-level data. + +NetworkX is for [the creation, manipulation, and study of complex networks](https://networkx.org/documentation/stable/index.html). [Pick the graph class first](https://networkx.org/documentation/stable/reference/introduction.html#graphs): Graph, DiGraph, MultiGraph, or MultiDiGraph. Store numeric edge data under [the `weight` key](https://networkx.org/documentation/stable/reference/introduction.html#nodes-and-edges), which algorithms like Dijkstra's shortest path read by default. Its drawing functions are basic, since its [main goal is graph analysis, not visualization](https://networkx.org/documentation/stable/reference/drawing.html), so export the graph to a dedicated visualization tool. + +Shapely wraps the GEOS library for [manipulation and analysis of planar geometric objects](https://shapely.readthedocs.io/en/stable/). For many geometries, put them in a NumPy array and call its [vectorized functions instead of a Python loop](https://shapely.readthedocs.io/en/stable/#usage). It doesn't read or write data files or transform coordinate systems. Operations [presume the features sit in the same Cartesian plane](https://shapely.readthedocs.io/en/stable/manual.html), so project your data to a plane before you measure it. + +Colour provides [algorithms and datasets for colour science](https://colour.readthedocs.io/en/latest/). Manim [generates animations of technical concepts from Python code](https://docs.manim.community/en/stable/). It comes in more than one edition, and the one listed here is [the community-maintained one its docs recommend](https://docs.manim.community/en/stable/faq/installation.html#which-version-should-i-use), especially for beginners. Each animation lives in [the `construct()` method of a `Scene` subclass](https://docs.manim.community/en/stable/tutorials/quickstart.html), with helper functions outside the class. + +Keep your data in NumPy arrays as it moves between most of these libraries. ObsPy already holds each trace's samples in one, and Colour [recommends them as input](https://colour.readthedocs.io/en/latest/basics.html). For results you can reproduce, seed your random numbers: [pass a seed to `default_rng()`](https://numpy.org/doc/stable/reference/random/index.html#random-quick-start) in NumPy, and [give your Mesa model a `seed` argument](https://mesa.readthedocs.io/latest/best-practices.html#randomization).