mirror of
https://github.com/vinta/awesome-python.git
synced 2026-10-02 08:23:10 +08:00
Merge branch 'feature/category-page-redesign'
This commit is contained in:
@@ -33,7 +33,7 @@ Run the `preview-verdicts` skill: it generates the interactive review page and d
|
|||||||
|
|
||||||
## 5. Execute
|
## 5. Execute
|
||||||
|
|
||||||
One commit per section: body lists each removal with its reason and downloads figure; restructures, tier moves, and reorders ride the same commit. Format-only outcomes (no removals) are a single style commit. `make test` before every commit, `make build` after the last one. Generic commit helpers tend to split a section audit into structural and per-subcategory commits — if that happens, squash back to one commit per section. Done when the tree is clean, tests passed before each commit, and the build count reconciles with the adjudicated changes.
|
One commit per section: body lists each removal with its reason and downloads figure; restructures, tier moves, and reorders ride the same commit. A restructure that renames or dissolves a category or subcategory adds `"old path": "new path"` to `website/data/redirects.json` in that commit, so the old URL keeps its search ranking; the build fails when a redirect target no longer exists, so a later rename also updates the entries pointing at the renamed path. Format-only outcomes (no removals) are a single style commit. `make test` before every commit, `make build` after the last one. Generic commit helpers tend to split a section audit into structural and per-subcategory commits — if that happens, squash back to one commit per section. Done when the tree is clean, tests passed before each commit, and the build count reconciles with the adjudicated changes.
|
||||||
|
|
||||||
## 6. Record
|
## 6. Record
|
||||||
|
|
||||||
|
|||||||
@@ -13,6 +13,8 @@ __pycache__/
|
|||||||
website/output/
|
website/output/
|
||||||
website/data/*
|
website/data/*
|
||||||
!website/data/pypi_name_overrides.json
|
!website/data/pypi_name_overrides.json
|
||||||
|
!website/data/redirects.json
|
||||||
|
!website/data/category_intros/
|
||||||
|
|
||||||
# agents
|
# agents
|
||||||
.playwright-cli/
|
.playwright-cli/
|
||||||
|
|||||||
@@ -9,7 +9,15 @@ An opinionated guide to the best Python frameworks, libraries, and tools.
|
|||||||
[CONTRIBUTING.md](CONTRIBUTING.md) holds the admission rules, quality requirements, rejection rules, entry format, and ordering. Apply it whenever adding or removing an entry — direct commits included, not only PR reviews.
|
[CONTRIBUTING.md](CONTRIBUTING.md) holds the admission rules, quality requirements, rejection rules, entry format, and ordering. Apply it whenever adding or removing an entry — direct commits included, not only PR reviews.
|
||||||
|
|
||||||
- Every keep/drop reason must be verified against current online data at decision time — download counts, repo activity and archived status, PyPI metadata, project docs. Judging tiers — obvious choice vs challenger — also requires WebSearch evidence (adoption trajectory, community sentiment), not download counts alone. Training-data recollections are not evidence; label anything unverifiable as a judgment call.
|
- Every keep/drop reason must be verified against current online data at decision time — download counts, repo activity and archived status, PyPI metadata, project docs. Judging tiers — obvious choice vs challenger — also requires WebSearch evidence (adoption trajectory, community sentiment), not download counts alone. Training-data recollections are not evidence; label anything unverifiable as a judgment call.
|
||||||
|
- Checks stay read-only: judge whether a listed project installs or works from its PyPI metadata (`requires_python`, classifiers, wheel tags) and issue tracker, never by installing or running it, since that runs untrusted code.
|
||||||
- A download count is not automatically independent demand. When one listed entry depends on another, check `requires_dist` on PyPI before citing the depended-on entry's count: mkdocs-material hard-depends on mkdocs, so mkdocs' figure exceeded mkdocs-material's by only about one percent and nearly all of it was mkdocs-material pulling it in.
|
- A download count is not automatically independent demand. When one listed entry depends on another, check `requires_dist` on PyPI before citing the depended-on entry's count: mkdocs-material hard-depends on mkdocs, so mkdocs' figure exceeded mkdocs-material's by only about one percent and nearly all of it was mkdocs-material pulling it in.
|
||||||
- One entry per commit when adding or deleting entries. Exceptions: a prune sweep is one commit per section, its body listing each removal with its reason; format, wording, or categorization changes may be bundled. Cross-section re-homes ride the originating audit's commit (both sides of the move in one diff).
|
- One entry per commit when adding or deleting entries. Exceptions: a prune sweep is one commit per section, its body listing each removal with its reason; format, wording, or categorization changes may be bundled. Cross-section re-homes ride the originating audit's commit (both sides of the move in one diff).
|
||||||
- Resources sections are not project entries: out of audit scope, and the website never parses them.
|
- Resources sections are not project entries: out of audit scope, and the website never parses them.
|
||||||
- Sponsor placement never influences which projects get listed — see [SPONSORSHIP.md](SPONSORSHIP.md).
|
- Sponsor placement never influences which projects get listed — see [SPONSORSHIP.md](SPONSORSHIP.md).
|
||||||
|
|
||||||
|
## Website
|
||||||
|
|
||||||
|
- Model the layout on https://www.placestoread.xyz: the whole list on one page, dense rows, a row that expands inline with its content aligned to the Name column, sorting by column headers, full-text search, tags that filter, a divider under the header row instead of a strong top border, a solid dark footer, and minimal decoration. Keep that model (no card grid, modal details, or pagination) unless the maintainer asks, and keep green out of the palette.
|
||||||
|
- `.section-shell` (`--shell-max: 84rem`) is the only width cap: sections, tables, rows, CTAs, and paragraphs stay uncapped, since the maintainer reads on wide screens and prose-width advice like 65-75ch doesn't apply. Remove inner `max-width` rules you come across, and ask with a concrete reason before adding one.
|
||||||
|
- Pick font sizes one step larger than feels right (`--text-base` over `--text-sm`, 1.75rem over 1.5rem for a heading), and shrink an existing size only when the maintainer asks: they asked for bigger type 11 times across 8 sessions and never for smaller.
|
||||||
|
- When you style one link, tag, or label, give its peers the same style in the same change: hero, nav, footer, project, and sponsor links share hover styles, and tag variants build on `.tag`.
|
||||||
|
|||||||
@@ -9,7 +9,15 @@ An opinionated guide to the best Python frameworks, libraries, and tools.
|
|||||||
[CONTRIBUTING.md](CONTRIBUTING.md) holds the admission rules, quality requirements, rejection rules, entry format, and ordering. Apply it whenever adding or removing an entry — direct commits included, not only PR reviews.
|
[CONTRIBUTING.md](CONTRIBUTING.md) holds the admission rules, quality requirements, rejection rules, entry format, and ordering. Apply it whenever adding or removing an entry — direct commits included, not only PR reviews.
|
||||||
|
|
||||||
- Every keep/drop reason must be verified against current online data at decision time — download counts, repo activity and archived status, PyPI metadata, project docs. Judging tiers — obvious choice vs challenger — also requires WebSearch evidence (adoption trajectory, community sentiment), not download counts alone. Training-data recollections are not evidence; label anything unverifiable as a judgment call.
|
- Every keep/drop reason must be verified against current online data at decision time — download counts, repo activity and archived status, PyPI metadata, project docs. Judging tiers — obvious choice vs challenger — also requires WebSearch evidence (adoption trajectory, community sentiment), not download counts alone. Training-data recollections are not evidence; label anything unverifiable as a judgment call.
|
||||||
|
- Checks stay read-only: judge whether a listed project installs or works from its PyPI metadata (`requires_python`, classifiers, wheel tags) and issue tracker, never by installing or running it, since that runs untrusted code.
|
||||||
- A download count is not automatically independent demand. When one listed entry depends on another, check `requires_dist` on PyPI before citing the depended-on entry's count: mkdocs-material hard-depends on mkdocs, so mkdocs' figure exceeded mkdocs-material's by only about one percent and nearly all of it was mkdocs-material pulling it in.
|
- A download count is not automatically independent demand. When one listed entry depends on another, check `requires_dist` on PyPI before citing the depended-on entry's count: mkdocs-material hard-depends on mkdocs, so mkdocs' figure exceeded mkdocs-material's by only about one percent and nearly all of it was mkdocs-material pulling it in.
|
||||||
- One entry per commit when adding or deleting entries. Exceptions: a prune sweep is one commit per section, its body listing each removal with its reason; format, wording, or categorization changes may be bundled. Cross-section re-homes ride the originating audit's commit (both sides of the move in one diff).
|
- One entry per commit when adding or deleting entries. Exceptions: a prune sweep is one commit per section, its body listing each removal with its reason; format, wording, or categorization changes may be bundled. Cross-section re-homes ride the originating audit's commit (both sides of the move in one diff).
|
||||||
- Resources sections are not project entries: out of audit scope, and the website never parses them.
|
- Resources sections are not project entries: out of audit scope, and the website never parses them.
|
||||||
- Sponsor placement never influences which projects get listed — see [SPONSORSHIP.md](SPONSORSHIP.md).
|
- Sponsor placement never influences which projects get listed — see [SPONSORSHIP.md](SPONSORSHIP.md).
|
||||||
|
|
||||||
|
## Website
|
||||||
|
|
||||||
|
- Model the layout on https://www.placestoread.xyz: the whole list on one page, dense rows, a row that expands inline with its content aligned to the Name column, sorting by column headers, full-text search, tags that filter, a divider under the header row instead of a strong top border, a solid dark footer, and minimal decoration. Keep that model (no card grid, modal details, or pagination) unless the maintainer asks, and keep green out of the palette.
|
||||||
|
- `.section-shell` (`--shell-max: 84rem`) is the only width cap: sections, tables, rows, CTAs, and paragraphs stay uncapped, since the maintainer reads on wide screens and prose-width advice like 65-75ch doesn't apply. Remove inner `max-width` rules you come across, and ask with a concrete reason before adding one.
|
||||||
|
- Pick font sizes one step larger than feels right (`--text-base` over `--text-sm`, 1.75rem over 1.5rem for a heading), and shrink an existing size only when the maintainer asks: they asked for bigger type 11 times across 8 sessions and never for smaller.
|
||||||
|
- When you style one link, tag, or label, give its peers the same style in the same change: hero, nav, footer, project, and sponsor links share hover styles, and tag variants build on `.tag`.
|
||||||
|
|||||||
@@ -140,7 +140,7 @@ Hard-won sizing rules (do not relax):
|
|||||||
Depth comes from **tonal layers**, not heavy shadows.
|
Depth comes from **tonal layers**, not heavy shadows.
|
||||||
|
|
||||||
- The page is a quiet warm canvas (`--bg-page`). The content shell is slightly brighter paper (`--bg-paper`). The sponsor band, CTA backgrounds, and inline decorative blocks step up to `--bg-paper-strong`.
|
- The page is a quiet warm canvas (`--bg-page`). The content shell is slightly brighter paper (`--bg-paper`). The sponsor band, CTA backgrounds, and inline decorative blocks step up to `--bg-paper-strong`.
|
||||||
- The hero is the one place that uses real atmosphere: subtle grid, slow sheen, warm radial gradients on a dark earthy ground (`--hero-bg-start` → `--hero-bg-mid` → `--hero-bg-end`). The sheen and any other motion respect `prefers-reduced-motion`.
|
- The hero and the category guide band are the only places with real atmosphere: warm gradients on a dark earthy ground (`--hero-bg-start` → `--hero-bg-mid` → `--hero-bg-end`). Only the hero adds the subtle grid and slow sheen. The sheen and any other motion respect `prefers-reduced-motion`.
|
||||||
- The footer is a single tonal block in `--footer-bg`, no internal gradients.
|
- The footer is a single tonal block in `--footer-bg`, no internal gradients.
|
||||||
- Two depth treatments are allowed and only these two. The search input combines a 1px inset highlight (`--search-inset`) with a soft warm drop shadow (`--search-shadow`, intensified by `--search-focus-shadow` on focus). The primary CTA button (`.hero-action-primary`) carries a warm drop shadow for press affordance. Both shadows are soft, warm-tinted, and tied to interactive elements. No new drop shadows on cards, panels, rows, or static decoration.
|
- Two depth treatments are allowed and only these two. The search input combines a 1px inset highlight (`--search-inset`) with a soft warm drop shadow (`--search-shadow`, intensified by `--search-focus-shadow` on focus). The primary CTA button (`.hero-action-primary`) carries a warm drop shadow for press affordance. Both shadows are soft, warm-tinted, and tied to interactive elements. No new drop shadows on cards, panels, rows, or static decoration.
|
||||||
- No glassmorphism as default decoration.
|
- No glassmorphism as default decoration.
|
||||||
@@ -160,8 +160,11 @@ The shape language is overwhelmingly **pill on small, zero radius on large**.
|
|||||||
The component vocabulary is small and table-led. Source of truth: `website/static/style.css`.
|
The component vocabulary is small and table-led. Source of truth: `website/static/style.css`.
|
||||||
|
|
||||||
- **Table-driven index** (the hero of the page). Sticky header, sortable columns, click-to-expand rows that indent under the Name column. Modeled on placestoread.xyz. Not a card grid.
|
- **Table-driven index** (the hero of the page). Sticky header, sortable columns, click-to-expand rows that indent under the Name column. Modeled on placestoread.xyz. Not a card grid.
|
||||||
|
- **Group rows**. Section, subcategory, and group pages list rows in README order. Section pages add one H2 row per use case; group pages add one per section, linking to that section's page. A note replaces the sort arrow until someone sorts a column, which flattens the list and hides the group rows.
|
||||||
- **Filter tags** (`.tag`). `--accent-soft` background with `--accent-deep` text. Pill shape. Hover swaps to `--highlight` background with `--tag-hover-border` border and ink text. Active state uses the warm `--tag-active-start` → `--tag-active-end` gradient with hero-ink text. Tag variants (`tag-group`, `tag-source`) inherit the base `.tag` style today and differ only at narrow widths (`tag-group` hides under 960px). Add a new variant only when a real visual difference is needed.
|
- **Filter tags** (`.tag`). `--accent-soft` background with `--accent-deep` text. Pill shape. Hover swaps to `--highlight` background with `--tag-hover-border` border and ink text. Active state uses the warm `--tag-active-start` → `--tag-active-end` gradient with hero-ink text. Tag variants (`tag-group`, `tag-source`) inherit the base `.tag` style today and differ only at narrow widths (`tag-group` hides under 960px). Add a new variant only when a real visual difference is needed.
|
||||||
- **Hero**. Magazine-cover headline, dark earthy ground, kicker and proof microcopy, primary CTA button using `--hero-btn-start` / `--hero-btn-end`. Subtle grid plus slow sheen. Respects `prefers-reduced-motion`.
|
- **Hero**. Magazine-cover headline, dark earthy ground, kicker and proof microcopy, primary CTA button using `--hero-btn-start` / `--hero-btn-end`. Subtle grid plus slow sheen. Respects `prefers-reduced-motion`.
|
||||||
|
- **Jump links**. One per use case in a category hero, plus one for the guide. Underlined like the intro's links, with a trailing ↓. Never tag-styled: tags filter the table or open another page, jump links only scroll.
|
||||||
|
- **Guide band**. Everything in a category intro after its "How to choose" list. It sits below the table on the hero's dark gradient, with its H2 above the text.
|
||||||
- **Sponsor band**. Sits in the README header on `--bg-paper-strong`. Editorial layout, not a logo wall. Sponsor links share the global accent treatment.
|
- **Sponsor band**. Sits in the README header on `--bg-paper-strong`. Editorial layout, not a logo wall. Sponsor links share the global accent treatment.
|
||||||
- **CTA**. Warm `--cta-bg`, full-bleed within shell. The button itself uses accent tokens.
|
- **CTA**. Warm `--cta-bg`, full-bleed within shell. The button itself uses accent tokens.
|
||||||
- **Footer**. Dark warm charcoal, part of the same system. Footer links share the global hover and focus treatment.
|
- **Footer**. Dark warm charcoal, part of the same system. Footer links share the global hover and focus treatment.
|
||||||
|
|||||||
+104
-11
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,17 @@
|
|||||||
|
Rather than hand-build a Python admin panel, use Flask-Admin in Flask apps. Django has one built in, which Unfold restyles for dashboards and Grappelli reskins.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A Flask app: Flask-Admin
|
||||||
|
- Dashboards and a Tailwind CSS look on the Django admin: Unfold
|
||||||
|
- A grid-based skin for the Django admin: Grappelli
|
||||||
|
|
||||||
|
Flask-Admin [solves the boring problem](https://flask-admin.readthedocs.io/en/stable/) of building an admin interface on top of an existing data model. Where the Django admin works from sensible defaults, Flask-Admin [leaves it to you](https://flask-admin.readthedocs.io/en/stable/advanced/#migrating-from-django) to tell it what to display and how. Register [a `ModelView` for each model](https://flask-admin.readthedocs.io/en/stable/introduction/#adding-model-views) with `admin.add_view()` to get its list, create, and edit views. To shape them, subclass `ModelView` and set attributes like `column_list` and `column_searchable_list`. It works with several ORMs and MongoDB, and if you don't know where to start, its docs point you to [SQLAlchemy](https://flask-admin.readthedocs.io/en/stable/advanced/#using-different-database-backends). For a page that isn't tied to a model, like analytics, extend [`BaseView`](https://flask-admin.readthedocs.io/en/stable/introduction/#standalone-views).
|
||||||
|
|
||||||
|
Unfold is [a modern Django admin theme](https://unfoldadmin.com/) for building dashboards, internal tools, and business applications with Tailwind CSS. Put `"unfold"` [first in `INSTALLED_APPS`](https://unfoldadmin.com/docs/installation/quickstart/), before `django.contrib.admin`. Your admin URLs stay the same. Your admin classes then inherit from `unfold.admin.ModelAdmin`, since the default class gives you unstyled forms without Unfold's features. That includes the User and Group admins Django registers for you, which you [unregister and register again](https://unfoldadmin.com/docs/installation/auth/) with Unfold's class. Build [the dashboard](https://unfoldadmin.com/docs/configuration/dashboard/) by overriding the admin index template, fed with data from a callback.
|
||||||
|
|
||||||
|
Grappelli is [a grid-based skin](https://django-grappelli.readthedocs.io/en/latest/) for the Django admin. Its FAQ says that [if you're pleased with how the original admin looks](https://django-grappelli.readthedocs.io/en/latest/faq.html#why-should-i-use-grappelli), you probably shouldn't use it. Put `grappelli` before `django.contrib.admin` in `INSTALLED_APPS`, and [include its URLs](https://django-grappelli.readthedocs.io/en/latest/quickstart.html#setup), which related lookups and autocompletes need. Its [customization examples](https://django-grappelli.readthedocs.io/en/latest/customization.html) build on Django's own `admin.ModelAdmin` and inline classes. They add options for collapsible fieldsets, drag-and-drop inline sorting, and autocomplete lookups.
|
||||||
|
|
||||||
|
An admin panel is for trusted users. Django's docs [limit the admin](https://docs.djangoproject.com/en/stable/ref/contrib/admin/#overview) to an organization's internal management tool, not your whole front end. By default, it lets in only users with `is_staff` set. Unfold and Grappelli both run on it, so the same holds for them. When you need a process-centric interface instead of one built around tables and fields, write your own views. Flask-Admin leaves [keeping unwanted users out](https://flask-admin.readthedocs.io/en/stable/introduction/#authorization-permissions) to you: override `is_accessible` on your admin views with your own login check.
|
||||||
|
|
||||||
|
On Django, admin add-ons built for the stock look may not match a theme. Before you switch, check [Unfold's integrations](https://unfoldadmin.com/docs/) or [Grappelli's third-party list](https://django-grappelli.readthedocs.io/en/latest/thirdparty.html) for the ones you rely on.
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
LangChain is the place to start among Python libraries for AI agents, and LangGraph gives you control of every step. vLLM serves your own models.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Skills for your coding agent: Django AI Skills, Sentry Skills, or Trail of Bits Skills
|
||||||
|
- A first agent: LangChain, or LangGraph to control every step
|
||||||
|
- An agent built on one vendor's platform: OpenAI Agents SDK or Claude Agent SDK
|
||||||
|
- A ready-made personal assistant: Hermes Agent, or AstrBot for chat apps
|
||||||
|
- Prompts tuned against a metric instead of by hand: DSPy
|
||||||
|
- Structured output, RAG, or agent memory: Instructor, LlamaIndex, or Mem0
|
||||||
|
- Running pre-trained models: Transformers
|
||||||
|
- Serving a model: vLLM, or MLX LM on Apple silicon
|
||||||
|
- One API for many LLM providers: LiteLLM
|
||||||
|
- Image and video generation: Diffusers
|
||||||
|
- Fine-tuning: PEFT, Unsloth, or Axolotl
|
||||||
|
- Speech: Whisper for speech to text, Kitten TTS for text to speech
|
||||||
|
|
||||||
|
New to agents? LangGraph's own docs [recommend LangChain's prebuilt agents](https://docs.langchain.com/oss/python/langgraph/overview), which run on LangGraph: give an agent a model, tools, and a prompt, and the loop is handled for you. Drop down to LangGraph for [needs that combine deterministic and agentic workflows](https://docs.langchain.com/oss/python/langchain/overview). You don't need LangChain to use LangGraph.
|
||||||
|
|
||||||
|
Pydantic AI is the pick when you want your type checker to cover the agent too. Give the agent an output type, and [every run comes back as a validated Pydantic model](https://pydantic.dev/docs/ai/overview/); when validation fails, the model is asked to try again. Tools and instructions get their dependencies through typed injection, so you can swap in a test double in unit tests.
|
||||||
|
|
||||||
|
CrewAI splits the work into Crews, teams of role-playing agents, and Flows, event-driven workflows that hold state. For production apps, its docs recommend [starting with a Flow](https://docs.crewai.com/en/concepts/production-architecture) and calling Crews from it.
|
||||||
|
|
||||||
|
OpenAI Agents SDK keeps [the primitives few](https://openai.github.io/openai-agents-python/): agents, handoffs, and guardrails, with tracing built in. It also runs [non-OpenAI models](https://openai.github.io/openai-agents-python/models/). Claude Agent SDK runs [Claude Code as a library](https://code.claude.com/docs/en/agent-sdk/overview): the same built-in tools, permissions, sessions, and hooks, inside your own process.
|
||||||
|
|
||||||
|
Instructor gets validated data out of an LLM into a Pydantic model, with retries when validation fails. Its own docs draw the line: [Instructor for extraction, Pydantic AI for agents](https://python.useinstructor.com/).
|
||||||
|
|
||||||
|
DSPy has you [write signatures, not prompts](https://dspy.ai/). Give it examples and a metric, and its optimizers tune the prompts for you.
|
||||||
|
|
||||||
|
LlamaIndex is a [data framework](https://github.com/run-llama/llama_index) for LLM apps: it loads, indexes, and queries your documents. Install `llama-index` to start, or `llama-index-core` plus only the integrations you need.
|
||||||
|
|
||||||
|
Mem0 adds [memory that persists across sessions](https://docs.mem0.ai/). Self-host the open-source version, or use the managed platform. OpenViking is [AGPL-licensed](https://github.com/volcengine/OpenViking/blob/main/LICENSE), where Mem0 is Apache-licensed.
|
||||||
|
|
||||||
|
Transformers runs pre-trained models from the Hugging Face Hub. Start with [`pipeline()`](https://huggingface.co/docs/transformers/pipeline_tutorial): pick a task and a model, and it handles preprocessing and output. Diffusers works the same way for [diffusion models](https://huggingface.co/docs/diffusers/index).
|
||||||
|
|
||||||
|
vLLM serves a model behind an [OpenAI-compatible API](https://docs.vllm.ai/en/latest/getting_started/quickstart.html). SGLang does too, and its [RadixAttention caches shared prefixes](https://docs.sglang.io/), which helps when requests share a long prompt. On Apple silicon, [MLX LM](https://github.com/ml-explore/mlx-lm) runs and fine-tunes models locally.
|
||||||
|
|
||||||
|
LiteLLM puts many LLM providers behind one OpenAI-style API. Use [the Python SDK in your code, or run the proxy as a gateway](https://docs.litellm.ai/docs/) when a platform team needs keys, budgets, and spend tracking across projects.
|
||||||
|
|
||||||
|
PEFT [trains a small set of extra parameters](https://huggingface.co/docs/peft/index) instead of the whole model, and works with Transformers and Diffusers. Unsloth's docs [recommend starting with QLoRA](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide). Axolotl drives [the whole pipeline from one YAML file](https://docs.axolotl.ai/): preprocessing, training, evaluation, quantization, and inference.
|
||||||
|
|
||||||
|
Whisper is a [general-purpose speech recognition model](https://github.com/openai/whisper) that also translates speech and identifies languages. Microsoft marks VibeVoice for [research and development only](https://github.com/microsoft/VibeVoice).
|
||||||
|
|
||||||
|
For text to speech, [Kitten TTS runs on CPU](https://github.com/KittenML/KittenTTS) without a GPU. gTTS calls [Google Translate's undocumented speech endpoint](https://github.com/pndurette/gTTS), so it needs the internet and can break without notice.
|
||||||
|
|
||||||
|
The skill repos aren't pip packages: they install into your coding agent, not your app. Django AI Skills and Sentry Skills follow the [Agent Skills](https://agentskills.io/) open format.
|
||||||
|
|
||||||
|
Write your app against the OpenAI API format, and you can switch between a hosted model and your own: vLLM, SGLang, and the LiteLLM proxy all speak it.
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
Read The Algorithms to see how algorithms work, not to ship. Install Sorted Containers when your code needs a Python algorithms library for sorted collections.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Sorted lists, dicts, and sets: Sorted Containers
|
||||||
|
- Learning how an algorithm works: The Algorithms
|
||||||
|
- Algorithm code to read that also installs with pip: algorithms
|
||||||
|
- A state machine bound to an object you already have: transitions
|
||||||
|
- Design patterns, and which ones Python doesn't need: python-patterns
|
||||||
|
- A state machine declared as a class, up to full statecharts: python-statemachine
|
||||||
|
|
||||||
|
Python's standard library is great [until you need a sorted collections type](https://grantjenks.com/docs/sortedcontainers/). Sorted Containers fills that gap in pure Python, with no C compiler to install. `SortedList` is its core type: it [keeps its values in ascending order](https://grantjenks.com/docs/sortedcontainers/introduction.html#sorted-list) as you add them with `add()` or `update()`. `SortedDict` is a dict that also keeps a sorted list of its keys, and `SortedSet` is a set that keeps a sorted list of its values.
|
||||||
|
|
||||||
|
The Algorithms implements [algorithms in Python for education](https://github.com/TheAlgorithms/Python), from sorts and searches to graphs and dynamic programming. Browse them by topic in its [directory](https://github.com/TheAlgorithms/Python/blob/master/DIRECTORY.md).
|
||||||
|
|
||||||
|
The algorithms package puts each data structure and algorithm [in a self-contained file](https://github.com/keon/algorithms) with docstrings, type hints, and complexity notes, written to be read and learned from. It also installs with pip, so your code can import what you've read, like `from algorithms.graph import dijkstra`.
|
||||||
|
|
||||||
|
transitions is [a lightweight, object-oriented state machine](https://github.com/pytransitions/transitions) that you bind to an object you already have. [Pass `Machine`](https://github.com/pytransitions/transitions#basic-initialization) your model, its states, and its transitions as dicts, each with a trigger, a source, and a destination. The model then gets a method for each trigger, like `evaporate()`.
|
||||||
|
|
||||||
|
python-patterns is [a collection of design patterns and idioms](https://github.com/faif/python-patterns), one file per pattern, grouped as creational, structural, behavioral, and more. Its README asks you to care more about why you pick a pattern than how you implement it. [Its anti-patterns section](https://github.com/faif/python-patterns#-anti-patterns) lists the ones not recommended in Python: modules are already singletons, so use a module-level variable instead of a Singleton class.
|
||||||
|
|
||||||
|
python-statemachine defines [flat state machines or full statecharts](https://python-statemachine.readthedocs.io/en/latest/) in a declarative class that works in both sync and async code. Statecharts add compound states, parallel regions, and history. States are class attributes, like `green = State(initial=True)`. `green.to(yellow)` declares a transition, and `|` combines transitions into one event, as in `cycle = green.to(yellow) | yellow.to(red)`.
|
||||||
|
|
||||||
|
The Algorithms says its implementations [may be less efficient than the standard library's](https://github.com/TheAlgorithms/Python), so where the standard library has an algorithm, use its version in your code. python-statemachine's docs show how to [rewrite a hand-written State pattern declaratively](https://python-statemachine.readthedocs.io/en/latest/how-to/coming_from_state_pattern.html), like the one in python-patterns' `state.py`. For more places to learn and practice algorithms, see [awesome-algorithms](https://github.com/tayllan/awesome-algorithms).
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
I/O-bound or CPU-bound? I/O-bound code wants a Python async library, and asyncio comes built in. CPU-bound work goes to a concurrent.futures process pool.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- I/O-bound code written with async/await: asyncio
|
||||||
|
- CPU-bound work, or blocking calls, in a pool of processes or threads: concurrent.futures
|
||||||
|
- Trio-style task groups and cancel scopes on asyncio, or a library that runs on both asyncio and Trio: AnyIO
|
||||||
|
- A faster, drop-in event loop for an asyncio app on Linux or macOS: uvloop
|
||||||
|
- Structured concurrency from the ground up, with its own libraries: Trio
|
||||||
|
- Existing synchronous code, made concurrent without async/await: gevent
|
||||||
|
- Network servers and clients with protocols built in, like SSH, mail, and DNS: Twisted
|
||||||
|
- Processes you manage yourself, talking through queues and pipes: multiprocessing
|
||||||
|
|
||||||
|
asyncio is the standard library's way to [write concurrent code with async/await](https://docs.python.org/3/library/asyncio.html), and many async web servers, database drivers, and task queues build on it. Start your program with `asyncio.run()`, [called once as the main entry point](https://docs.python.org/3/library/asyncio-runner.html#asyncio.run). App code should [rarely need the event loop object](https://docs.python.org/3/library/asyncio-eventloop.html) itself. Run related tasks in an `asyncio.TaskGroup`: when one task fails, it [cancels the rest, which `gather()` doesn't](https://docs.python.org/3/library/asyncio-task.html#asyncio.gather). CPU-heavy code holds up every task on the loop, so [run it in another process](https://docs.python.org/3/library/asyncio-dev.html#asyncio-handle-blocking): hand it to a `ProcessPoolExecutor` with `loop.run_in_executor()`.
|
||||||
|
|
||||||
|
concurrent.futures runs callables on threads or processes behind [the same interface](https://docs.python.org/3/library/concurrent.futures.html). `ThreadPoolExecutor` is [for overlapping I/O](https://docs.python.org/3/library/concurrent.futures.html#threadpoolexecutor). For CPU-bound work on a multi-core machine, the threading docs [advise processes](https://docs.python.org/3/library/threading.html) instead, and `ProcessPoolExecutor` runs them for you. It [takes only picklable functions and arguments](https://docs.python.org/3/library/concurrent.futures.html#processpoolexecutor), though. Use either executor [in a `with` block](https://docs.python.org/3/library/concurrent.futures.html#concurrent.futures.Executor.shutdown), which shuts it down and waits for its work to finish.
|
||||||
|
|
||||||
|
AnyIO brings [Trio-like structured concurrency to asyncio](https://anyio.readthedocs.io/en/stable/). Code written against its API runs unmodified on asyncio or Trio, so a library built on it doesn't choose for its users. Its docs see [strong merits in its APIs for applications too](https://anyio.readthedocs.io/en/stable/why.html), starting with cancel scopes for [more predictable cancellation](https://anyio.readthedocs.io/en/stable/why.html#design-problems-with-cancellation). Start with `anyio.run(main)`, which [runs on asyncio unless you pass `backend="trio"`](https://anyio.readthedocs.io/en/stable/basics.html#running-async-programs). Spawn tasks in a [task group](https://anyio.readthedocs.io/en/stable/tasks.html): when one child task raises, the rest are cancelled.
|
||||||
|
|
||||||
|
uvloop is [a drop-in replacement for asyncio's event loop](https://github.com/MagicStack/uvloop), built on libuv. Its README prefers `uvloop.run(main())`, which configures `asyncio.run()` to use uvloop, so the rest of your asyncio code stays the same. uvloop runs on Linux and macOS.
|
||||||
|
|
||||||
|
Trio has [an obsessive focus on usability and correctness](https://trio.readthedocs.io/en/stable/). Child tasks run in a nursery, opened with `async with trio.open_nursery()`, and [Trio never discards their exceptions](https://trio.readthedocs.io/en/stable/tutorial.html#okay-let-s-see-something-cool-already). Functions [take no timeout arguments](https://trio.readthedocs.io/en/stable/reference-core.html#blocking-and-non-blocking-methods): you wrap the code in a cancel scope like `trio.move_on_after()`. Trio runs its own event loop, so asyncio functions [don't work inside `trio.run()`](https://trio.readthedocs.io/en/stable/tutorial.html#task-switching-illustrated). Check [the Trio library list](https://trio.readthedocs.io/en/stable/awesome-trio-libraries.html) for what you need first.
|
||||||
|
|
||||||
|
gevent uses greenlets to give you [a synchronous API on top of an event loop](https://www.gevent.org/intro.html). Its monkey patching swaps the standard library's blocking sockets for cooperative ones, so code that knows nothing about gevent runs concurrently. Most programs should [patch everything with `monkey.patch_all()`](https://www.gevent.org/intro.html#beyond-sockets), and the main module should do it [before any other imports](https://www.gevent.org/api/gevent.monkey.html).
|
||||||
|
|
||||||
|
Twisted is [an event-based framework for internet applications](https://github.com/twisted/twisted) that ships clients and servers for HTTP, SSH, IMAP, POP3, SMTP, DNS, IRC, and XMPP. Write new Twisted code as `async def` coroutines, which its docs [prefer over `inlineCallbacks`](https://docs.twisted.org/en/stable/core/howto/defer-intro.html#inline-callbacks-using-yield), and start one with [`Deferred.fromCoroutine()`](https://docs.twisted.org/en/stable/core/howto/defer-intro.html#coroutines-with-async-await).
|
||||||
|
|
||||||
|
multiprocessing runs work in subprocesses, so it can [use every processor on a machine](https://docs.python.org/3/library/multiprocessing.html#introduction). Its own docs point to `ProcessPoolExecutor` as the higher-level interface for pooled tasks. Reach for multiprocessing when you need what it adds, like killing a running process or passing data through queues and pipes. Its guidelines say to [avoid shared state](https://docs.python.org/3/library/multiprocessing.html#all-start-methods) and keep the data moving between processes small.
|
||||||
|
|
||||||
|
On an event loop, a blocking call holds up every other task, so push it to a worker thread: asyncio has [`asyncio.to_thread()`](https://docs.python.org/3/library/asyncio-task.html#asyncio.to_thread), AnyIO [`to_thread.run_sync()`](https://anyio.readthedocs.io/en/stable/threads.html), and Trio [`trio.to_thread.run_sync()`](https://trio.readthedocs.io/en/stable/reference-core.html#threads-if-you-must). When multiprocessing or `ProcessPoolExecutor` runs your code in child processes, define its functions in a module. Guard your entry point with `if __name__ == '__main__':` too, so each new process can [import your main module safely](https://docs.python.org/3/library/multiprocessing.html#multiprocessing-safe-main-import).
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
Each job has its own pick: librosa is the Python audio processing library for music analysis, MoviePy edits video from a script, and Mutagen tags audio files.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Cutting, joining, and fading audio files: pydub
|
||||||
|
- Music and audio analysis: librosa
|
||||||
|
- Editing video or making GIFs from a script: MoviePy
|
||||||
|
- Real-time video from cameras and network streams: VidGear
|
||||||
|
- Reading and writing tags across formats: Mutagen
|
||||||
|
- Reading tags only: tinytag
|
||||||
|
- Tagging and organizing your music collection: beets
|
||||||
|
|
||||||
|
pydub gives you a simple, high-level interface to cut, join, and fade audio. It opens and saves WAV files in pure Python, but [other formats like MP3 need FFmpeg](https://github.com/jiaaro/pydub#dependencies), so install FFmpeg with it. An AudioSegment is [immutable](https://github.com/jiaaro/pydub#quickstart): every operation returns a new one, so you can chain them, and every length and position is in milliseconds.
|
||||||
|
|
||||||
|
librosa gives you [the foundational algorithms and tools for music information retrieval](https://librosa.org/doc/latest/index.html). By default, `librosa.load` resamples the signal and mixes stereo down to mono, and those defaults [suit most analysis tasks](https://librosa.org/doc/latest/auto_tutorials/01-intro/01-load.html); pass `sr=None` to keep the file's own sampling rate. When you need more control than `load` gives you, such as writing files, [its docs recommend using its audio I/O backend directly](https://librosa.org/doc/latest/ioformats.html).
|
||||||
|
|
||||||
|
MoviePy is for [automating video editing](https://zulko.github.io/moviepy/getting_started/quick_presentation.html): processing many videos, composing them in complicated ways, or making videos and GIFs on a web server. A script loads clips, modifies them, puts them together, and writes the result. Modifying a clip [returns a new clip and leaves the original alone](https://zulko.github.io/moviepy/user_guide/modifying.html), and the computation happens at the final render. Open file clips in a `with` block, or [call `close()`](https://zulko.github.io/moviepy/user_guide/loading.html) when you're done, since each one holds a subprocess and a lock on the file. MoviePy can't stream video. For frame-by-frame analysis, its docs send you to a computer vision library.
|
||||||
|
|
||||||
|
VidGear is a framework for [real-time media applications](https://abhitronix.github.io/vidgear/latest/) built on OpenCV and FFmpeg. All its APIs [keep OpenCV's coding syntax](https://abhitronix.github.io/vidgear/latest/switch_from_cv/). Each task has [its own gear](https://abhitronix.github.io/vidgear/latest/gears/): CamGear reads cameras, network streams, and streaming sites in multiple threads. WriteGear writes frames to a video file or network stream, and StreamGear transcodes video into adaptive streaming formats. [Install OpenCV first](https://abhitronix.github.io/vidgear/latest/installation/pip_install/), since the core functions need it.
|
||||||
|
|
||||||
|
Mutagen reads and writes tags with [roughly the same API across all tag formats](https://mutagen.readthedocs.io/en/latest/). `mutagen.File` [guesses the file type](https://mutagen.readthedocs.io/en/latest/user/gettingstarted.html). ID3 tags in MP3 files are highly structured; for common keys, use [the simpler EasyID3 interface](https://mutagen.readthedocs.io/en/latest/user/id3.html). Mutagen is GPL-licensed; if you only read tags, MIT-licensed tinytag avoids that.
|
||||||
|
|
||||||
|
tinytag only reads metadata, and [writing support will not be added](https://github.com/tinytag/tinytag): its README points you to Mutagen for that. It's pure Python with no dependencies and gives you the same API for every format. `TinyTag.get()` returns an object with attributes like `artist` and `duration`.
|
||||||
|
|
||||||
|
beets is a command-line music library manager, not a library you import: it [catalogs your collection and improves its metadata](https://beets.io/) as it goes. Install it [as a standalone tool](https://beets.readthedocs.io/en/stable/guides/installation.html), isolated from your system Python and other packages. `beet import` can modify and move your files, so [back up first and import a few albums at a time](https://beets.readthedocs.io/en/stable/guides/main.html). [Plugins](https://beets.readthedocs.io/en/stable/plugins/index.html) add commands, fetch extra data during import, and add metadata sources.
|
||||||
|
|
||||||
|
FFmpeg sits under most of these projects: pydub needs it for any format other than WAV, MoviePy runs on it, and VidGear's WriteGear and StreamGear wrap it. When you only want to convert a video file or turn images into a movie, [call FFmpeg directly](https://zulko.github.io/moviepy/getting_started/quick_presentation.html). MoviePy's own docs say it's faster and uses less memory than going through MoviePy.
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
Letting users log in with Google takes one Python authentication library: django-allauth on Django, which does password sign-up too, or Authlib elsewhere.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Sign-up, password login, and Google or GitHub login on a Django site: django-allauth
|
||||||
|
- Google or GitHub login, OAuth API calls, or an OAuth server outside Django: Authlib
|
||||||
|
- Adding OAuth to a web framework or HTTP library you maintain: oauthlib
|
||||||
|
- Your own OAuth 2.0 server on Django, or tokens for a Django REST framework API: django-oauth-toolkit
|
||||||
|
- Signing and verifying JSON Web Tokens: PyJWT
|
||||||
|
- Permissions you grant per object to users and groups: django-guardian
|
||||||
|
- Permissions that follow from the object's data, with no extra tables: django-rules
|
||||||
|
|
||||||
|
django-allauth exists to handle [local and social accounts in one app](https://docs.allauth.org/en/latest/introduction/index.html): a user can sign up with a password or log in with Google. Follow the [quickstart](https://docs.allauth.org/en/latest/installation/quickstart.html): add its authentication backend, apps, and middleware, then include `allauth.urls` under `accounts/`. Its login, logout, and password views can replace the ones in `django.contrib.auth.urls`. Give each OAuth provider its client ID and secret in a [`SocialApp` or the `SOCIALACCOUNT_PROVIDERS` setting](https://docs.allauth.org/en/latest/installation/quickstart.html#post-installation). For a single-page or mobile app, add [`allauth.headless`](https://docs.allauth.org/en/latest/headless/introduction.html).
|
||||||
|
|
||||||
|
Authlib covers both sides of OAuth 2.0 and OpenID Connect. As a client, it comes in [two kinds](https://docs.authlib.org/en/latest/oauth2/client/index.html): HTTP clients for scripts and service-to-service calls, and web clients that log users in through Flask, Django, Starlette, or FastAPI. For login, [create an `OAuth` registry and `register()` each provider](https://docs.authlib.org/en/latest/oauth2/client/web/index.html). For an OpenID Connect provider like Google, pass its discovery URL as [`server_metadata_url`](https://docs.authlib.org/en/latest/oauth2/client/web/index.html#parsing-id-token), and Authlib reads the other endpoints from it. To [run your own OAuth 2.0 or OpenID Connect provider](https://docs.authlib.org/en/latest/oauth2/authorization-server/index.html), it has server integrations for Flask and Django.
|
||||||
|
|
||||||
|
PyJWT [encodes and decodes JSON Web Tokens](https://pyjwt.readthedocs.io/en/latest/). Treat every token as untrusted input: [hard-code the `algorithms` you pass to `jwt.decode()`](https://pyjwt.readthedocs.io/en/latest/algorithms.html#specifying-an-algorithm), never read them from the token's own header, and don't mix HS and RS algorithms. For tokens from an OpenID Connect provider, point [`PyJWKClient` at its JWKS endpoint](https://pyjwt.readthedocs.io/en/latest/usage.html#retrieve-rsa-signing-keys-from-a-jwks-endpoint) instead of hard-coding its public keys.
|
||||||
|
|
||||||
|
oauthlib implements OAuth's logic [without assuming a web framework or HTTP request object](https://github.com/oauthlib/oauthlib), so a framework or library maintainer can write a thin layer on top and get OAuth support. Its FAQ says [most people use it indirectly](https://oauthlib.readthedocs.io/en/latest/faq.html#how-do-i-use-oauthlib-with-google-twitter-and-other-providers): django-oauth-toolkit and django-allauth both build on it. To build a provider on it, most of the work is [a `RequestValidator`](https://oauthlib.readthedocs.io/en/latest/oauth2/server.html#implement-a-validator) that maps its checks to your storage.
|
||||||
|
|
||||||
|
django-oauth-toolkit is an [OAuth 2.0 authorization server for Django](https://django-oauth-toolkit.readthedocs.io/en/latest/): it issues and manages tokens from your existing project, and can also protect a Django or Django REST framework API. Django REST framework's docs [recommend it for OAuth 2.0](https://www.django-rest-framework.org/api-guide/authentication/#django-oauth-toolkit). [Install it](https://django-oauth-toolkit.readthedocs.io/en/latest/install.html) by adding `oauth2_provider` to `INSTALLED_APPS`, including its URLs under `o/`, and migrating. For an API, [set `OAuth2Authentication`](https://django-oauth-toolkit.readthedocs.io/en/latest/rest-framework/getting_started.html) as Django REST framework's authentication class and guard views with `TokenHasScope`.
|
||||||
|
|
||||||
|
Django has [a foundation for object permissions but no implementation](https://docs.djangoproject.com/en/stable/topics/auth/customizing/#handling-object-permissions), and django-guardian fills it with an [extra authentication backend](https://django-guardian.readthedocs.io/en/latest/configuration/). Grant a permission on one object with `assign_perm()`, then check it with Django's own [`user.has_perm(perm, obj)`](https://django-guardian.readthedocs.io/en/latest/userguide/checks/). Django REST framework's `DjangoObjectPermissions` [names it as an example backend](https://www.django-rest-framework.org/api-guide/permissions/#djangoobjectpermissions).
|
||||||
|
|
||||||
|
django-rules gives Django object permissions [without a database](https://github.com/dfunckt/django-rules). A permission is a rule built from predicates, plain functions like `is_book_author(user, book)` that you combine with `|`, `&`, and `~`. Add its [backend to `AUTHENTICATION_BACKENDS`](https://github.com/dfunckt/django-rules#configuring-django), register each permission with `rules.add_perm()`, and check it with `user.has_perm()`. Keep predicates and rules in their own modules, and let [`AutodiscoverRulesConfig` import each app's `rules.py`](https://github.com/dfunckt/django-rules#best-practices) at startup.
|
||||||
|
|
||||||
|
On either side of OAuth, use PKCE. django-oauth-toolkit's security guide maps [the OAuth security recommendations of RFC 9700](https://django-oauth-toolkit.readthedocs.io/en/latest/security.html) to its settings: PKCE is one, and the implicit and password grants must not be used. Authlib's client [turns PKCE on with `code_challenge_method`](https://docs.authlib.org/en/latest/oauth2/client/web/index.html#oauth-2-0-code-challenge). Also, don't pick JWTs for sessions just because they sound stateless. django-allauth's docs point out that [if a token must stop working at logout, you need state to revoke it](https://docs.allauth.org/en/latest/headless/token-strategies/jwt-tokens.html). JWTs help most when each service checks tokens without asking a central server.
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
Where Make would go, a Python build tool takes over, and SCons compiles C and C++. Invoke runs shell commands as tasks, and doit reruns only what changed.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- C, C++, or Fortran code compiled from source: SCons
|
||||||
|
- Shell commands run as named tasks with CLI flags: Invoke
|
||||||
|
- Tasks that rerun only when their input files change: doit
|
||||||
|
- Your own CLI program built from tasks: Invoke
|
||||||
|
|
||||||
|
SCons builds software from source code. Its site calls it [an improved, cross-platform substitute for the classic Make utility](https://scons.org/), with dependency analysis for C, C++, and Fortran built in. Your build file [is a Python script](https://scons.org/doc/production/HTML/scons-user/ch02s05.html#sect-sconstruct-python) named `SConstruct`: [put `Program('hello.c')` in it](https://scons.org/doc/production/HTML/scons-user/ch02.html) and run `scons`. You declare what to build, and SCons works out the build order, [whatever order you call the builders in](https://scons.org/doc/production/HTML/scons-user/ch02s05.html#sect-order-independent). For a source tree with subdirectories, [split the build into `SConscript` files](https://scons.org/doc/production/HTML/scons-user/ch14.html#sect-sconscript-files) that the top-level `SConstruct` pulls in. SCons [puts a correct build first](https://scons.org/doc/production/HTML/scons-user/pr01.html#sect-principles): by default, it [decides a file has changed from a hash of its contents](https://scons.org/doc/production/HTML/scons-user/ch06.html#sect-contentsigs), not its modification time.
|
||||||
|
|
||||||
|
Invoke turns your project's shell commands into Python tasks. [Write them in a `tasks.py`](https://docs.pyinvoke.org/en/stable/getting-started.html#defining-and-running-task-functions) as `@task` functions whose first argument is a context, and [run commands with `c.run()`](https://docs.pyinvoke.org/en/stable/getting-started.html#running-shell-commands). [Each parameter becomes a CLI flag](https://docs.pyinvoke.org/en/stable/getting-started.html#task-parameters), and `invoke --list` shows your tasks. A task can name [pre-tasks](https://docs.pyinvoke.org/en/stable/getting-started.html#declaring-pre-tasks) to run first, like `clean` before `build`. When one flat list of tasks gets crowded, [group them into namespaces](https://docs.pyinvoke.org/en/stable/getting-started.html#creating-namespaces) with `Collection`.
|
||||||
|
|
||||||
|
Invoke can also [power your own CLI program](https://docs.pyinvoke.org/en/stable/concepts/library.html#reusing-invoke-s-cli-module-as-a-distinct-binary), with your tasks as its commands. It [sticks to local commands](https://www.pyinvoke.org/faq.html#why-was-invoke-split-off-from-the-fabric-project) and leaves servers and network commands to a separate library built on it.
|
||||||
|
|
||||||
|
doit runs only what changed. [Tasks live in a `dodo.py`](https://pydoit.org/tasks.html#intro): each function whose name starts with `task_` returns a dict, and its [actions](https://pydoit.org/tasks.html#actions) are shell commands or Python functions. [List a task's `file_dep` and `targets`](https://pydoit.org/tasks.html#dependencies-targets), and doit skips it when its dependencies haven't changed and its targets exist. Inputs [don't have to be files](https://pydoit.org/dependencies.html#uptodate): an `uptodate` check such as `config_changed` reruns a task when a config string or dict changes.
|
||||||
|
|
||||||
|
Looking for the tool that builds your package's wheels and publishes them to PyPI? That's packaging, covered in [Package Management](/categories/package-management/).
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
Skip hand-written dunder methods, and when you compare Python dataclass alternatives, pick attrs for its validators and converters. Box gives dicts dot access.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Classes with generated dunder methods, validators, and converters: attrs
|
||||||
|
- Dicts read with dot notation, nested ones included: python-box
|
||||||
|
- Looking up a key by its value, with both directions in sync: bidict
|
||||||
|
- A faster drop-in for the `uuid` module: uuid-utils
|
||||||
|
|
||||||
|
attrs [writes the dunder methods](https://www.attrs.org/en/latest/) that implement object protocols, so you don't have to. For new code, its docs [recommend the modern API](https://www.attrs.org/en/latest/names.html#tl-dr): `@define` on the class and `field()` for an attribute, or `@frozen` for immutable instances, with slots on by default. Type annotations stay optional. The standard library's dataclasses [gave up features for simplicity](https://www.attrs.org/en/latest/why.html#data-classes), like validators and converters, so a class that needs them goes to attrs. It [isn't a full serialization library](https://www.attrs.org/en/latest/overview.html#what-attrs-is-not), though: to serialize and validate your attrs classes, its docs point you to its sibling project, cattrs.
|
||||||
|
|
||||||
|
python-box installs Box, a [`dict` subclass](https://github.com/cdgriffith/Box) whose keys you can also read as attributes, even ones like `"imdb stars"`. Nested dicts and lists become Box and BoxList objects, so dot access works all the way down: `movie_box.Robin_Hood_Men_in_Tights.imdb_stars`. It [reads and writes JSON, YAML, TOML, and msgpack](https://github.com/cdgriffith/Box/wiki/Converters) with methods like `from_json()` and `to_json()`, and `to_dict()` gives you plain dicts back. Keep Box for data whose keys change: attrs' docs say a dict with a fixed and known set of keys [is an object, not a hash](https://www.attrs.org/en/latest/why.html#dicts), and belongs in a class.
|
||||||
|
|
||||||
|
bidict gives you [bidirectional mappings](https://bidict.readthedocs.io/intro.html) that work like dicts: look up a value by its key as usual, or a key by its value through `.inverse`, which stays in sync as you update the mapping. A single dict holding both directions mixes keys with values. Modeling the mapping correctly takes two one-way mappings kept in sync, [which is what bidict does](https://bidict.readthedocs.io/intro.html#why-can-t-i-just-use-a-dict) under the hood.
|
||||||
|
|
||||||
|
uuid-utils is a [fast, drop-in replacement for Python's `uuid` module](https://aminalaee.github.io/uuid-utils/latest/), powered by Rust. Import it under the module's name, `import uuid_utils as uuid`, and call `uuid.uuid4()` or `uuid.uuid7()` as usual. Django and some other frameworks require the standard library's own `uuid.UUID` instances. For those, [import `uuid_utils.compat`](https://aminalaee.github.io/uuid-utils/latest/#compatibility-with-python-uuid) instead, which returns them and still outperforms the standard library.
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
In memory, on disk, or on a cache server: the Python caching library you want is cachetools, DiskCache, or dogpile.cache. For HTTP responses, it's Hishel.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Function results in one process's memory, with a size limit or a time-to-live: cachetools
|
||||||
|
- A cache on local disk that processes on one machine share, with no server to run: DiskCache
|
||||||
|
- A Django cache backend on local disk: DiskCache
|
||||||
|
- One cache on memcached or Redis for many processes or servers: dogpile.cache
|
||||||
|
- HTTP responses in HTTPX or Requests: Hishel
|
||||||
|
- `Cache-Control` headers and response caching in FastAPI or another ASGI app: Hishel
|
||||||
|
- Django querysets, invalidated when a model changes: Cacheops
|
||||||
|
|
||||||
|
cachetools offers [variants of the standard library's `@lru_cache`](https://cachetools.readthedocs.io/en/stable/) with more cache algorithms, including a `TTLCache` whose items [expire after a time-to-live](https://cachetools.readthedocs.io/en/stable/#cachetools.TTLCache). Wrap a function in [`@cached`](https://cachetools.readthedocs.io/en/stable/#cachetools.cached) and pass it the cache to use. The cache classes [aren't thread-safe](https://cachetools.readthedocs.io/en/stable/#cache-implementations), so when threads share a cache, give `@cached` a `threading.Lock`. The cache lives in your process's memory, so each worker process of a web app keeps its own copy.
|
||||||
|
|
||||||
|
DiskCache is a [disk and file backed cache library](https://grantjenks.com/docs/diskcache/) in pure Python, built on SQLite, with no other process to run. Create a `Cache` with a directory path: [two `Cache` objects on the same directory](https://grantjenks.com/docs/diskcache/tutorial.html#cache) can live in separate processes, so the workers on one machine share one cache. Wrap a function in [`@cache.memoize()`](https://grantjenks.com/docs/diskcache/tutorial.html#fanoutcache), which takes arguments like `lru_cache`'s. In Django, set [`diskcache.DjangoCache`](https://grantjenks.com/docs/diskcache/tutorial.html#djangocache) as the cache `BACKEND`. Keep the directory on a local disk, since SQLite [isn't recommended on NFS mounts](https://grantjenks.com/docs/diskcache/tutorial.html#caveats).
|
||||||
|
|
||||||
|
dogpile.cache is [a caching API over backends of any variety](https://dogpilecache.sqlalchemy.org/en/latest/), memcached and Redis among them. You ask it for a value and hand it a function that [creates the value only when needed](https://dogpilecache.sqlalchemy.org/en/latest/usage.html#overview). When the value expires, one worker regenerates it instead of every worker at once. Create a region with `make_region()` at import time, decorate functions with `@region.cache_on_arguments()`, and [call `configure()` later](https://dogpilecache.sqlalchemy.org/en/latest/usage.html#rudimentary-usage) with the backend and expiration time from your config file. When several processes share one Redis or memcached server, turn on the backend's [`distributed_lock`](https://dogpilecache.sqlalchemy.org/en/latest/api.html#dogpile.cache.backends.redis.RedisBackend) so that lock covers them all.
|
||||||
|
|
||||||
|
Hishel caches HTTP responses by [RFC 9111](https://hishel.com/overview.html), the HTTP caching rules browsers follow, in both sync and async code. With HTTPX, [swap your client for Hishel's cache-enabled one](https://hishel.com/httpx.html#quick-start), or [give a client you already have its cache transport](https://hishel.com/httpx.html#cache-transports). With Requests, [mount its cache adapter](https://hishel.com/requests.html#quick-start) on a `Session`. In FastAPI, [a dependency sets `Cache-Control` headers](https://hishel.com/fastapi.html#quick-start) for browsers and CDNs, and its ASGI middleware also caches responses on your server.
|
||||||
|
|
||||||
|
Cacheops caches Django querysets in Redis and [invalidates them](https://github.com/Suor/django-cacheops#user-content-invalidation) on `save()`, `delete()`, and many-to-many changes. [Add it to `INSTALLED_APPS`](https://github.com/Suor/django-cacheops#user-content-setup) and point it at its own Redis database, as its docs highly recommend. Turn caching on per app with `'app_name.*'`, since `'*.*'` can also cache tables you don't mean to, like migrations. Call `.cache()` on a queryset to cache it by hand, and wrap a function in [`@cached_as(Article)`](https://github.com/Suor/django-cacheops#user-content-usage) to drop its result whenever an `Article` changes.
|
||||||
|
|
||||||
|
Plan how stale entries leave the cache, too. cachetools' decorators add a [`cache_clear()`](https://cachetools.readthedocs.io/en/stable/#memoizing-decorators) function, DiskCache [evicts every key with a given tag](https://grantjenks.com/docs/diskcache/tutorial.html#cache), and a function decorated by dogpile.cache takes [`invalidate()`](https://dogpilecache.sqlalchemy.org/en/latest/api.html#dogpile.cache.region.CacheRegion.cache_on_arguments) with the same arguments you'd call it with.
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
Python's own argparse covers basic command-line apps. Beyond that, pick Click or Typer as your Python CLI library, and style the output with Rich.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Basic command-line app with no dependencies: argparse
|
||||||
|
- Nested commands, or subcommands loaded lazily: Click
|
||||||
|
- Options declared once as type-hinted function parameters: Typer
|
||||||
|
- Interactive prompts and REPLs: prompt_toolkit
|
||||||
|
- CLI from existing code without writing a parser: Python Fire
|
||||||
|
- Progress bar for a loop: tqdm, or alive-progress for animated bars
|
||||||
|
- Colored text, tables, and logs: Rich, or colorama for ANSI colors on Windows
|
||||||
|
- Full-screen app in the terminal or a browser: Textual
|
||||||
|
- Terminal widgets on an event loop you already run: Urwid
|
||||||
|
- Full-screen forms and ASCII animations: asciimatics
|
||||||
|
|
||||||
|
argparse is in the standard library. Python's docs call it [the recommended choice](https://docs.python.org/3/library/optparse.html#choosing-an-argument-parser) when you have no more specific needs, since it gives you the most out of the box for the least code. It [writes the help and usage messages](https://docs.python.org/3/library/argparse.html) and reports invalid arguments for you. For commands written as decorated functions, the same docs point to Click, and for a CLI that works with static type checking, to Typer.
|
||||||
|
|
||||||
|
Click builds a CLI from [commands declared with decorators](https://click.palletsprojects.com/en/stable/quickstart/). You can nest them to any depth, and Click can [load subcommands lazily](https://click.palletsprojects.com/en/stable/) at runtime. It also [reads option values from environment variables](https://click.palletsprojects.com/en/stable/why/) and comes with helpers for ANSI colors, terminal size, and launching editors.
|
||||||
|
|
||||||
|
Typer builds a CLI from your function signatures: you [declare each argument and option once](https://typer.tiangolo.com/), as a type-hinted function parameter. Write those hints with `Annotated`, which its tutorial [recommends wherever possible](https://typer.tiangolo.com/tutorial/arguments/optional/). Its docs say you get [the best results by pairing it with Rich](https://typer.tiangolo.com/tutorial/printing/): Typer structures the commands and options, and Rich displays the output.
|
||||||
|
|
||||||
|
prompt_toolkit is for interactive input. It can [replace GNU readline](https://python-prompt-toolkit.readthedocs.io/en/stable/) or build full-screen apps. For a REPL, use a [`PromptSession`](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/asking_for_input.html), which keeps the history for the whole session.
|
||||||
|
|
||||||
|
Python Fire turns any Python object into a CLI: [call `fire.Fire()`](https://github.com/google/python-fire/blob/master/docs/guide.md) at the end of your program. Its README pitches it for [developing and debugging your code](https://github.com/google/python-fire), exploring existing code, and turning other people's code into a CLI. It takes each argument's type from the value you pass, not from the function signature.
|
||||||
|
|
||||||
|
tqdm adds a progress bar to any loop: [wrap the iterable in `tqdm()`](https://github.com/tqdm/tqdm). Import it from `tqdm.auto`, which picks the console bar or the Jupyter widget for you. alive-progress wraps the loop in an [`alive_bar` context manager](https://github.com/rsalmei/alive-progress) instead, with a spinner that speeds up and slows down with your throughput.
|
||||||
|
|
||||||
|
Rich writes colored text, tables, Markdown, and syntax-highlighted code to the terminal. Start with its [drop-in `print`](https://rich.readthedocs.io/en/stable/introduction.html), then create [one `Console` at module level](https://rich.readthedocs.io/en/stable/console.html) for the rest of your app. For colored logs, send the logging module's output through [`RichHandler`](https://rich.readthedocs.io/en/stable/logging.html).
|
||||||
|
|
||||||
|
colorama makes ANSI escape codes work on Windows and does nothing on other platforms. If that's all you need, call [`just_fix_windows_console()`](https://github.com/tartley/colorama).
|
||||||
|
|
||||||
|
Textual apps run in the terminal, and the `textual serve` command from textual-dev [serves them in a browser](https://textual.textualize.io/guide/devtools/). Build one by [subclassing `App`](https://textual.textualize.io/guide/app/), and style its widgets with [CSS](https://textual.textualize.io/guide/CSS/).
|
||||||
|
|
||||||
|
Urwid is a [console widget construction set](https://urwid.org/manual/overview.html) rather than a finished UI library. It runs on [your choice of event loop](https://github.com/urwid/urwid): asyncio, or another one you already use. It's [licensed under the LGPL](https://github.com/urwid/urwid/blob/master/COPYING), while Textual and asciimatics use permissive licenses.
|
||||||
|
|
||||||
|
asciimatics does [full-screen text UIs, from interactive forms to ASCII animations](https://github.com/peterbrittain/asciimatics): create a Screen, build a Scene from Effect objects, and let the Screen play it.
|
||||||
|
|
||||||
|
Ship your CLI as an installable package with a `[project.scripts]` entry point, not a file users run with `python`. [Click recommends it](https://click.palletsprojects.com/en/stable/entry-points/), and shell completion in Click and [Typer](https://typer.tiangolo.com/tutorial/package/) works through that entry point. Test your commands in-process: Click and Typer both provide a [`CliRunner`](https://click.palletsprojects.com/en/stable/testing/), and Textual's [`run_test()`](https://textual.textualize.io/guide/testing/) drives your app as if you were using the keyboard and mouse.
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
Swap cURL for HTTPie, and your database's own client for a Python CLI tool with autocompletion: pgcli, mycli, litecli, or IRedis.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Calling HTTP APIs from the terminal: HTTPie
|
||||||
|
- A database shell with autocompletion: pgcli for PostgreSQL, mycli for MySQL, litecli for SQLite, IRedis for Redis
|
||||||
|
- Downloading video or audio from YouTube and other sites: yt-dlp
|
||||||
|
- A new project from a template: Cookiecutter, or Copier to pull in later template changes
|
||||||
|
- A shell you script in Python: xonsh
|
||||||
|
- A tmux session with its windows and panes, from one YAML file: tmuxp
|
||||||
|
|
||||||
|
HTTPie is a command-line HTTP client for [testing, debugging, and generally interacting with APIs](https://httpie.io/docs/cli) and HTTP servers. It formats and colors the output. A request reads like `http PUT pie.dev/put X-API-Token:123 name=John`: `Header:Value` items set headers, and `field=value` items fill the body, which HTTPie [sends as JSON by default](https://httpie.io/docs/cli/json).
|
||||||
|
|
||||||
|
pgcli, mycli, litecli, and IRedis are terminal clients with autocompletion and syntax highlighting, one per database. Give [pgcli](https://github.com/dbcli/pgcli) a database name or a `postgresql://` URI, and [litecli](https://github.com/dbcli/litecli) the path to a SQLite file. mycli also [works with MariaDB](https://github.com/dbcli/mycli), Percona, TiDB, and Apache Doris. IRedis behaves like Redis's own client in most cases, but it's [safer on production servers](https://github.com/laixintao/iredis): it stops you from accidentally running dangerous commands like `KEYS *`.
|
||||||
|
|
||||||
|
yt-dlp downloads audio and video from [thousands of sites](https://github.com/yt-dlp/yt-dlp). Install ffmpeg too: yt-dlp [needs it to merge separate video and audio files](https://github.com/yt-dlp/yt-dlp#dependencies).
|
||||||
|
|
||||||
|
Cookiecutter creates projects from templates, and its templates [work for any language](https://github.com/cookiecutter/cookiecutter), not only Python. [Run `cookiecutter gh:audreyfeldroy/cookiecutter-pypackage`](https://cookiecutter.readthedocs.io/en/stable/usage.html), answer its prompts, and you get a new project from that GitHub template.
|
||||||
|
|
||||||
|
Copier also generates projects from templates, but it calls itself [a code lifecycle management tool](https://copier.readthedocs.io/en/stable/comparisons/). When the template changes, `copier update` brings the changes into projects you already generated. [It works best](https://copier.readthedocs.io/en/stable/updating/) when both are in Git: the template tagged, your project clean, with the `.copier-answers.yml` file that records your answers. Templates can run code: Cookiecutter's [pre- and post-generate scripts](https://github.com/cookiecutter/cookiecutter), Copier's tasks. [Generate projects only from templates you trust](https://copier.readthedocs.io/en/stable/generating/).
|
||||||
|
|
||||||
|
xonsh is a shell whose language is [a superset of Python](https://github.com/xonsh/xonsh), with shell commands built in. It isn't POSIX-compatible, so don't make it your login shell with `chsh`. Its docs [recommend a xonsh profile in your terminal emulator](https://xon.sh/install.html#before-installing) instead.
|
||||||
|
|
||||||
|
tmuxp launches a whole tmux session [from one YAML or JSON file](https://tmuxp.git-pull.com/quickstart/): its windows, its panes, and the commands in them. Save the file as `.tmuxp.yaml` in a project, and [`tmuxp load path/to/project/`](https://github.com/tmux-python/tmuxp) builds the session.
|
||||||
|
|
||||||
|
Several of these tools' docs have you [install them as CLI tools](https://github.com/cookiecutter/cookiecutter), and add them to a project only to use them from Python. Looking for a library to build your own Python CLI tool? That's [CLI Development](/categories/cli-development/).
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
Both Python CMS picks run on Django: Wagtail for page types your developers define, like articles, and django CMS for editors building pages on the live site.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Content types your developers define, like articles and events: Wagtail
|
||||||
|
- Editors composing pages from reusable components on the live site: django CMS
|
||||||
|
|
||||||
|
Wagtail is [not an instant website in a box](https://docs.wagtail.org/en/stable/getting_started/the_zen_of_wagtail.html#wagtail-is-not-an-instant-website-in-a-box): expect to write code. Start a project with [`wagtail start`](https://docs.wagtail.org/en/stable/getting_started/quick_install.html). Each page type is [a Django model that inherits from `Page`](https://docs.wagtail.org/en/stable/topics/pages.html), so give each kind of content its own type with its own fields. An event page with a date and a location [can show up in a calendar](https://docs.wagtail.org/en/stable/getting_started/the_zen_of_wagtail.html#a-cms-should-get-information-out-of-an-editor-s-head-and-into-a-database-as-efficiently-and-directly-as-possible), and a styled heading on a generic page can't. Use [StreamField](https://docs.wagtail.org/en/stable/topics/streamfield.html) for pages without a fixed structure, like blog posts, and [snippets](https://docs.wagtail.org/en/stable/topics/snippets/index.html) for content that doesn't need its own page, like headers and footers. For a headless site, use its [built-in API](https://docs.wagtail.org/en/stable/advanced_topics/api/index.html).
|
||||||
|
|
||||||
|
In django CMS, [editors compose pages from plugins](https://docs.django-cms.org/en/stable/explanation/philosophy.html#three-disciplines-three-surfaces) in a toolbar on the live site, and developers write standard Django. Each template [declares its placeholders](https://docs.django-cms.org/en/stable/tutorials/02-templates-placeholders.html) with `{% placeholder %}`, and `CMS_TEMPLATES` lists the templates editors can pick from. For a region that is the same on every page, like a footer, use [`{% static_alias %}`](https://docs.django-cms.org/en/stable/tutorials/02-templates-placeholders.html#a-reusable-region-with-static-alias), so the content is stored once. For content from your own app, ask [where it lives](https://docs.django-cms.org/en/stable/explanation/composition.html). If it fits in a placeholder on a page, write a plugin. If it has its own list view, detail view, and URL, mount it as an app with an apphook. The core [publishes as you edit](https://docs.django-cms.org/en/stable/explanation/publishing.html), which is rarely enough for a production site with editors, so add a versioning package.
|
||||||
|
|
||||||
|
Both are built on Django. Wagtail [deploys like a Django site](https://docs.wagtail.org/en/stable/deployment/index.html). You can add either one to an existing Django project: Wagtail [integrates into one](https://docs.wagtail.org/en/stable/getting_started/integrating_into_django.html), and django CMS [doesn't make you rebuild around it](https://docs.django-cms.org/en/stable/explanation/philosophy.html#implications-for-projects). Neither fits a very small or static site. For a two-page brochure, django CMS calls itself [overkill](https://docs.django-cms.org/en/stable/explanation/philosophy.html#when-django-cms-may-not-be-the-right-fit), and a static site generator is lighter.
|
||||||
@@ -0,0 +1,43 @@
|
|||||||
|
Run Ruff for Python code analysis: it lints and formats. Add a type checker, since Ruff doesn't check types, and pre-commit to run Ruff on every commit.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Rules on which modules may import which: Import Linter
|
||||||
|
- Dead code: Vulture, or repowise to index the repo for your agent
|
||||||
|
- Deeply nested functions that are hard to read: complexipy
|
||||||
|
- Several analysis tools behind one command: Prospector
|
||||||
|
- Checks before every commit: pre-commit
|
||||||
|
- Linting, formatting, and import sorting: Ruff
|
||||||
|
- Formatting without Ruff: Black, with isort on its black profile
|
||||||
|
- Deeper inference and checks of your own: Pylint
|
||||||
|
- Flake8 plugins Ruff doesn't have: Flake8
|
||||||
|
- Security issues: Bandit
|
||||||
|
- Refactoring: Rope
|
||||||
|
- Type checking: mypy, or Pyright, ty, or Pyrefly to also check unannotated code
|
||||||
|
- Type hints from the types your code sees at runtime: MonkeyType
|
||||||
|
|
||||||
|
Ruff lints and formats in one tool, and its FAQ lists what it [can replace](https://docs.astral.sh/ruff/faq/#which-tools-does-ruff-replace): Flake8 and dozens of its plugins, Black, and isort. To sort imports and format, [run the linter, then the formatter](https://docs.astral.sh/ruff/formatter/#sorting-imports): `ruff check --select I --fix`, then `ruff format`. When you turn on a new rule in an existing codebase, [`--add-noqa`](https://docs.astral.sh/ruff/tutorial/#adding-rules) marks the current violations, so the rule only applies to new code.
|
||||||
|
|
||||||
|
Pick one formatter and stay with it. Ruff's formatter is designed as a drop-in replacement for Black, but it's [not meant to be used interchangeably with Black](https://docs.astral.sh/ruff/formatter/) over time. If you pick Black, set isort's black profile [in a config file at the root of your repo](https://isort.readthedocs.io/en/latest/configuration/black_compatibility.html), so it applies however isort runs. To adopt Black, [reformat everything in one commit](https://black.readthedocs.io/en/stable/guides/introducing_black_to_your_project.html) and list that commit in `.git-blame-ignore-revs`, so `git blame` skips it.
|
||||||
|
|
||||||
|
Ruff can be a [drop-in replacement for Flake8](https://docs.astral.sh/ruff/faq/#how-does-ruffs-linter-compare-to-flake8) when your code is formatted with Black and uses few or no Flake8 plugins. Keep Flake8 when you depend on a plugin Ruff doesn't cover. In pre-commit, install the plugin [through `additional_dependencies`](https://flake8.pycqa.org/en/latest/user/using-hooks.html).
|
||||||
|
|
||||||
|
Pylint [doesn't trust your type hints](https://pylint.readthedocs.io/en/stable/). It infers the actual values instead, which makes it slower but finds more issues in code that isn't fully typed. You can also write plugins for checks of your own. On a legacy project, start with `--errors-only`, then turn more messages on over time. Because of its speed, Pylint's docs suggest running it [in CI or a pre-push hook](https://pylint.readthedocs.io/en/stable/user_guide/installation/pre-commit-integration.html), not on every commit.
|
||||||
|
|
||||||
|
Bandit [finds common security issues](https://github.com/PyCQA/bandit) by walking each file's syntax tree. On an existing project, [save a baseline](https://bandit.readthedocs.io/en/latest/start.html) to ignore the findings you've judged to be non-issues. When you skip a line, write `# nosec B602` instead of a bare `# nosec`, so [a new issue on that line still shows up](https://bandit.readthedocs.io/en/latest/config.html).
|
||||||
|
|
||||||
|
Ruff is [a linter, not a type checker](https://docs.astral.sh/ruff/faq/#how-does-ruff-compare-to-mypy-or-pyright-or-pyre), so add one. mypy is [designed for gradual typing](https://mypy.readthedocs.io/en/stable/), and by default it skips functions without annotations. On an existing codebase, its docs suggest [starting with part of the code](https://mypy.readthedocs.io/en/stable/existing_code.html), running it in CI early, turning on `check_untyped_defs` as soon as you can, and aiming for `mypy --strict`.
|
||||||
|
|
||||||
|
Pyright, ty, and Pyrefly each come with a language server for your editor, and they check unannotated code too. Pyright [checks all code regardless of annotations](https://github.com/microsoft/pyright/blob/main/docs/mypy-comparison.md). Commit its config, [run it in CI](https://github.com/microsoft/pyright/blob/main/docs/getting-started.md), and turn on strict mode file by file with `# pyright: strict`.
|
||||||
|
|
||||||
|
ty is [designed for adoption](https://docs.astral.sh/ty/), with support for partially typed code. Add it as a [dev dependency](https://docs.astral.sh/ty/installation/), so everyone runs the same version. Pick Pyrefly when your code uses [Pydantic, Django, or attrs](https://pyrefly.org/en/docs/compare/) and you don't want the overhead of plugins. To switch to it, `pyrefly init` [migrates your existing type checker config](https://pyrefly.org/en/docs/installation/), and `pyrefly suppress` marks the current errors as ignored, so you start from a clean check.
|
||||||
|
|
||||||
|
Import Linter checks [contracts on your imports](https://import-linter.readthedocs.io/en/stable/), such as layers: higher layers may import lower ones, not the other way around. Vulture finds unused code. For its false positives, it recommends [a whitelist over `noqa` comments](https://github.com/jendrikseipp/vulture). After you delete dead code, run it again, since it may find more. complexipy measures how hard code is for people to understand, and its docs say to [run it alongside Ruff](https://complexipy.com/comparison-with-ruff/): Ruff catches wide functions, complexipy catches deep ones. On a large existing codebase, [take a snapshot](https://complexipy.com/usage-guide/) first, so only new complexity fails.
|
||||||
|
|
||||||
|
Prospector wraps several analysis tools and aims to be [useful out of the box](https://github.com/prospector-dev/prospector). repowise indexes your repo for coding agents and developers, with dead code and git history in the same index. It's licensed under the [AGPL, or a commercial license](https://github.com/repowise-dev/repowise).
|
||||||
|
|
||||||
|
Rope is a refactoring library, and its wiki suggests [starting with its language server plugin](https://github.com/python-rope/rope/wiki/How-to-use-Rope-in-my-IDE-or-Text-editor%3F) before the native editor integrations. MonkeyType records the types your code sees at runtime and writes annotations from them, which its docs call [an informative first draft](https://monkeytype.readthedocs.io/en/latest/) for you to check and fix. mypy's docs suggest [collecting types from test runs](https://mypy.readthedocs.io/en/stable/existing_code.html) this way.
|
||||||
|
|
||||||
|
pre-commit runs your hooks [before every commit](https://pre-commit.com/). Run `pre-commit install` after every clone, `pre-commit run --all-files` when you add a hook, and the same command in CI. With Ruff's hooks, [put the lint hook with `--fix` before the formatter](https://docs.astral.sh/ruff/integrations/#pre-commit), since its fixes can leave code that needs reformatting.
|
||||||
|
|
||||||
|
The checkers here do static code analysis: they check your code without running it. MonkeyType is the exception, since it records types while your code runs. Run Ruff [in your editor](https://docs.astral.sh/ruff/editors/) and on every commit, and Pylint and your type checker in CI.
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
Start with OpenCV when you need a Python computer vision library for images and video. To train and run detection models, use Ultralytics YOLO.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Image and video processing: OpenCV
|
||||||
|
- Detection, segmentation, and pose models: Ultralytics YOLO
|
||||||
|
- Vision ops inside a PyTorch model: Kornia
|
||||||
|
- Dataset curation and model evaluation: FiftyOne
|
||||||
|
- OCR on clean, printed documents: pytesseract
|
||||||
|
- OCR on text in photos: EasyOCR
|
||||||
|
|
||||||
|
OpenCV comes as [four pip packages that share the `cv2` namespace](https://github.com/opencv/opencv-python), so install only one: opencv-python for the main modules, or opencv-contrib-python to add the extra modules. If you never call `cv2.imshow` or you build your GUI with another toolkit, install the headless variant of either one, which also makes Docker images smaller.
|
||||||
|
|
||||||
|
Ultralytics YOLO covers the [whole life of a model](https://docs.ultralytics.com/modes/): train, validate, predict, export, and track, from Python or the `yolo` command. Its docs recommend [starting training from a pretrained model](https://docs.ultralytics.com/modes/train/). To deploy, [export it](https://docs.ultralytics.com/modes/export/) to ONNX, TensorRT, CoreML, or another format. The code and the models you train with it are [AGPL-3.0](https://www.ultralytics.com/license), so unless you open-source your whole project, you need an Enterprise License.
|
||||||
|
|
||||||
|
Kornia is a [differentiable computer vision library like OpenCV, with strong GPU support](https://kornia.readthedocs.io/en/latest/get-started/introduction.html). Every operator works on PyTorch tensors and supports autograd, so vision ops can run on the GPU and sit inside your training loop.
|
||||||
|
|
||||||
|
FiftyOne works on the data side of a model. [Load your dataset and your model's predictions into it](https://docs.voxel51.com/user_guide/basics.html), then see where the model succeeds and fails, and find mistakes in your labels. It [integrates with Ultralytics](https://docs.voxel51.com/integrations/ultralytics.html), so you can run and fine-tune YOLO models on FiftyOne datasets.
|
||||||
|
|
||||||
|
pytesseract [wraps the Tesseract OCR engine](https://github.com/madmaze/pytesseract), which you install on its own, then put on your PATH or point `tesseract_cmd` at. Tesseract suits clean, printed text and needs no GPU. To get better results, [improve the image first](https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html): Tesseract works best at 300 DPI or more, and retraining rarely helps unless you use an unusual font or a new language.
|
||||||
|
|
||||||
|
EasyOCR is a [general OCR that reads text in photos as well as in documents](https://www.jaided.ai/easyocr), in dozens of languages. It handles text in photos, where Tesseract struggles, but it's slow without a GPU. Create a `Reader` for your languages [once](https://github.com/JaidedAI/EasyOCR) and reuse it for every image, since that call loads the model into memory.
|
||||||
|
|
||||||
|
Ultralytics YOLO, Kornia, and EasyOCR run on PyTorch. When you need a specific CUDA build, install PyTorch before the library, as [Ultralytics](https://docs.ultralytics.com/quickstart/) and [Kornia](https://kornia.readthedocs.io/en/latest/get-started/installation.html) recommend. EasyOCR's README says the same for Windows.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
Settings that change between deployments belong in the environment, whatever Python configuration library you use. pydantic-settings reads them as typed fields.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Typed, validated settings from environment variables and `.env` files: pydantic-settings
|
||||||
|
- An INI file your users edit, with nothing to install: configparser
|
||||||
|
- A `.env` file loaded into the environment during development: python-dotenv
|
||||||
|
- Research code that composes its config and overrides it from the command line: Hydra
|
||||||
|
- Layered settings per environment, or settings for Django and Flask: Dynaconf
|
||||||
|
|
||||||
|
pydantic-settings loads [a settings class from environment variables or secrets files](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/). Subclass `BaseSettings` and declare each setting as a type-hinted field. [Any field you don't pass to the initializer gets its value from the environment](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/#usage), or its default when the variable isn't set. To read a `.env` file too, [set `env_file` in `model_config`](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/#dotenv-env-support), and its values get validated like the rest.
|
||||||
|
|
||||||
|
configparser comes with Python. It reads [a basic configuration language structured like Windows INI files](https://docs.python.org/3/library/configparser.html), so end users can customize your program by editing a file. Read values with mapping access, `config['section']['option']`, which the docs [prefer for new projects](https://docs.python.org/3/library/configparser.html#legacy-api-examples) over the legacy get and set methods. Values always come back as strings, so convert them with [`getint()`, `getfloat()`, and `getboolean()`](https://docs.python.org/3/library/configparser.html#supported-datatypes).
|
||||||
|
|
||||||
|
python-dotenv is for apps that [take their configuration from environment variables](https://saurabh-kumar.com/python-dotenv/#getting-started). In development, setting each one yourself isn't practical, so call `load_dotenv()` before the rest of your code. It reads a `.env` file when there is one and adds its values to `os.environ`, so your code reads them with `os.getenv()` as if they came from the real environment.
|
||||||
|
|
||||||
|
Hydra is a framework for [research and other complex applications](https://hydra.cc/docs/1.3/intro/). It builds a hierarchical configuration by composition, and you override it from config files and the command line. Decorate your entry point with [`@hydra.main()`](https://hydra.cc/docs/1.3/intro/#basic-example) and point it at a YAML config. Then change a value with an argument like `db.user=root`. For options you switch between, like two databases, [create a config group](https://hydra.cc/docs/1.3/intro/#composition-example) with one file per option, and pick one in the config's `defaults` list. With [`--multirun`](https://hydra.cc/docs/1.3/intro/#multirun), one command runs your function once for each configuration you list.
|
||||||
|
|
||||||
|
Dynaconf reads settings from [files in several formats, environment variables, and Vault or Redis](https://www.dynaconf.com/#features). Start with [`dynaconf init -f toml`](https://www.dynaconf.com/#using-python-only), which creates `config.py`, `settings.toml`, and `.secrets.toml`. Then import `settings` from `config` in your code. TOML is its default and most recommended format. To give development and production their own sections in one file, [set `environments=True`](https://www.dynaconf.com/#layered-environments-on-files). In a Django app, `django.conf.settings` [becomes a Dynaconf settings object](https://www.dynaconf.com/#using-django), and in Flask, `app.config` does.
|
||||||
|
|
||||||
|
Keep settings that change between deployments out of your code. The [twelve-factor app stores config in environment variables](https://12factor.net/config), and python-dotenv and Dynaconf both cite it. python-dotenv, pydantic-settings, and Dynaconf let the real environment win over files. `load_dotenv()` [doesn't override variables already set](https://saurabh-kumar.com/python-dotenv/#getting-started), pydantic-settings [ranks environment variables above `.env` and secrets files](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/#field-value-priority), and Dynaconf [prioritizes environment variables over files](https://www.dynaconf.com/envvars/). Keep secrets out of Git, too: add `.env` to your `.gitignore`, and `dynaconf init` adds `.secrets.*` there for you.
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
Don't touch a Python cryptography library's raw ciphers: encrypt with cryptography's Fernet or PyNaCl. Paramiko runs SSH, and ItsDangerous signs tokens.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Encrypting data, X.509 certificates, and other standard formats: cryptography
|
||||||
|
- Encryption and signatures between your own apps, with algorithms picked for you: PyNaCl
|
||||||
|
- Hashing passwords for storage: PyNaCl
|
||||||
|
- SSH clients and servers, and SFTP: Paramiko
|
||||||
|
- Signed links, tokens, and cookies that users may read but not change: ItsDangerous
|
||||||
|
|
||||||
|
cryptography has 2 layers: recipes that need few decisions, and low-level primitives it calls the "hazmat" layer. Its docs [recommend the recipes layer whenever possible](https://cryptography.io/en/stable/#layout), and hazmat only when necessary. For data you encrypt with a key, the recipe is [Fernet](https://cryptography.io/en/stable/fernet/): a message encrypted with it can't be read or changed without the key. To encrypt with a password, run it through a key derivation function first. The docs [recommend Argon2id](https://cryptography.io/en/stable/fernet/#using-passwords-with-fernet), and you keep the salt to derive the same key again. When Fernet doesn't fit, the docs point you to [authenticated encryption](https://cryptography.io/en/stable/hazmat/primitives/symmetric-encryption/) before any raw cipher, since encryption alone keeps data secret but doesn't stop tampering. The recipes also cover [X.509 certificates](https://cryptography.io/en/stable/x509/tutorial/), from signing requests to self-signed certs.
|
||||||
|
|
||||||
|
PyNaCl is a binding to libsodium, a fork of NaCl. cryptography's FAQ explains the split: cryptography is [general purpose and interoperable with existing systems](https://cryptography.io/en/stable/faq/#how-does-cryptography-compare-to-nacl-networking-and-cryptography-library), while NaCl gives you a set of hand-selected algorithms. If you prefer NaCl's design, the FAQ recommends PyNaCl. With a shared key, encrypt through `SecretBox` or `Aead` and let PyNaCl [generate a random nonce](https://pynacl.readthedocs.io/en/latest/secret/#nacl.secret.Aead.encrypt) for each message, which its docs strongly recommend. Between 2 parties, `Box` [authenticates both sides](https://pynacl.readthedocs.io/en/latest/public/#nacl-public-box), and `SealedBox` sends a message only the recipient can decrypt, without proving who sent it. To store passwords, `nacl.pwhash.str()` [hashes with argon2id by default](https://pynacl.readthedocs.io/en/latest/password_hashing/#password-storage-and-verification), and `nacl.pwhash.verify()` checks a password against the hash.
|
||||||
|
|
||||||
|
Paramiko implements SSHv2 in pure Python, as both client and server. For common client jobs like running remote commands or transferring files, its homepage [recommends a higher-level library built on it](https://www.paramiko.org/). Use Paramiko directly for low-level primitives or to run an SSH server in Python. Its client API [starts with `SSHClient`](https://docs.paramiko.org/en/latest/), and checking the server's host key is your job: call `load_system_host_keys()` before `connect()`. A host key it can't find is [rejected by default](https://docs.paramiko.org/en/latest/api/client.html#paramiko.client.SSHClient.set_missing_host_key_policy) with an `SSHException`. `open_sftp()` opens an SFTP session on the same connection.
|
||||||
|
|
||||||
|
ItsDangerous signs data, so you can send it somewhere untrusted and get it back: the receiver [can see the data but can't modify it](https://itsdangerous.palletsprojects.com/en/stable/) without your key. Flask's default sessions are [cookies signed with it](https://flask.palletsprojects.com/en/stable/api/#flask.sessions.SecureCookieSessionInterface). Signing hides nothing, so when the data must stay secret, encrypt it with Fernet instead. Its docs say you'll [typically want a serializer, not a signer](https://itsdangerous.palletsprojects.com/en/stable/concepts/#serializer-vs-signer). `URLSafeSerializer` turns data into a URL-safe string, and `URLSafeTimedSerializer` rejects tokens older than the `max_age` you pass to `loads()`. Give each use its own [salt](https://itsdangerous.palletsprojects.com/en/stable/concepts/#the-salt), so an activation link's signature won't pass as an upgrade link's.
|
||||||
|
|
||||||
|
Whichever pick you use, keep secret keys out of your code. ItsDangerous's docs say the key [shouldn't be saved in source code or committed to version control](https://itsdangerous.palletsprojects.com/en/stable/concepts/#the-secret-key), and suggest reading it from an environment variable. When you generate a key or token yourself, use `os.urandom()`, [never the `random` module](https://cryptography.io/en/stable/random-numbers/), which isn't cryptographically secure. Plan for rotation, too: cryptography's [`MultiFernet`](https://cryptography.io/en/stable/fernet/#cryptography.fernet.MultiFernet) and ItsDangerous both [take a list of keys](https://itsdangerous.palletsprojects.com/en/stable/concepts/#key-rotation), so old tokens keep working after you add a new key.
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
Once your data nears the size of your RAM, switch Python data analysis libraries from pandas to Polars. Once it lives in a database, Ibis runs your code there.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Data that fits in memory: pandas
|
||||||
|
- Handing DataFrames to libraries that expect pandas: pandas
|
||||||
|
- Large data on one machine, even bigger than your RAM: Polars
|
||||||
|
- Data already in a database or warehouse: Ibis
|
||||||
|
- Code you prototype locally, then run on a warehouse or cluster: Ibis
|
||||||
|
|
||||||
|
pandas aims to be [the fundamental high-level building block](https://pandas.pydata.org/docs/getting_started/overview.html) for practical data analysis in Python. Its DataFrame fits time series and tables with mixed column types, like an SQL table or a spreadsheet. It's built on NumPy to work with the rest of the scientific Python stack. In production code, the docs [recommend `.loc` and `.iloc` over `[]`](https://pandas.pydata.org/docs/user_guide/indexing.html) to select data. Set values in [one `.loc` call](https://pandas.pydata.org/docs/user_guide/copy_on_write.html#chained-assignment), not chained indexing like `df["foo"][mask] = 100`. pandas keeps everything in memory, so a dataset that takes a sizable share of your RAM [gets unwieldy](https://pandas.pydata.org/docs/user_guide/scale.html), and the docs point you to other libraries.
|
||||||
|
|
||||||
|
Polars is [built for multithreaded computing on a single machine](https://docs.pola.rs/user-guide/misc/comparison/#pandas), and its stricter API leads to fewer schema bugs. It has no index: rows are known by their position, so [no index state can change what a query means](https://docs.pola.rs/user-guide/migration/pandas/#polars-does-not-have-a-multi-indexindex). Make the lazy API [your default](https://docs.pola.rs/user-guide/migration/pandas/#be-lazy): start from a function like `scan_csv()` or call `.lazy()`, then `.collect()` the result. That lets the optimizer [filter rows and pick columns while reading the data](https://docs.pola.rs/user-guide/concepts/lazy-api/). Stay eager for exploratory work, when [you don't know yet what your query will look like](https://docs.pola.rs/user-guide/concepts/lazy-api/#when-to-use-which). Write expressions inside `select`, `with_columns`, `filter`, and `group_by`, since Polars code that [looks like pandas code](https://docs.pola.rs/user-guide/migration/pandas/#key-syntax-differences) likely runs slower than it should. For data that doesn't fit in memory, run the query on the [streaming engine](https://docs.pola.rs/user-guide/concepts/streaming/), which works through it in batches.
|
||||||
|
|
||||||
|
Ibis [compiles one Python dataframe API](https://ibis-project.org/why#how-does-ibis-work) into each backend's native language, mostly SQL, so the database or engine does the work. pandas' ecosystem page lists it for [bridging local Python and remote databases](https://pandas.pydata.org/community/ecosystem.html#ibis). Start on the default DuckDB backend, as the tutorial [recommends](https://ibis-project.org/tutorials/basics#install-ibis), then [change the connection string](https://ibis-project.org/why#scaling-up-and-out) to run the same code on PySpark, BigQuery, or Trino. Expressions are lazy: nothing runs until you call a method like `to_pandas()`, and only then does Ibis [send the compiled query](https://ibis-project.org/tutorials/coming-from/pandas) to the backend.
|
||||||
|
|
||||||
|
You can mix all three. Polars converts a DataFrame [with `to_pandas()`](https://docs.pola.rs/api/python/stable/reference/dataframe/api/polars.DataFrame.to_pandas.html), Ibis returns results [as pandas or Polars DataFrames](https://ibis-project.org/reference/expression-tables#ibis.expr.types.relations.Table.to_pandas), and pandas [can use PyArrow](https://pandas.pydata.org/docs/user_guide/pyarrow.html) to trade data with Arrow-based libraries like Polars. Do the heavy work in Polars or Ibis, and hand a pandas DataFrame to the libraries that expect one.
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
Loading data from APIs and databases into a warehouse with a Python ETL library is dlt's job. Fetching stock prices for your own research is yfinance's.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Data from APIs and databases into a warehouse or data lake: dlt
|
||||||
|
- Stock prices and financials for your own research: yfinance
|
||||||
|
- pandas DataFrames in and out of AWS services: AWS SDK for pandas (awswrangler)
|
||||||
|
- One pipeline for both batch data and live streams: Pathway
|
||||||
|
- Chinese market data: AKShare
|
||||||
|
- SEC filings and XBRL financial statements: EdgarTools
|
||||||
|
- Many data providers behind one API: OpenBB
|
||||||
|
|
||||||
|
dlt [loads data from messy sources into well-structured datasets](https://dlthub.com/docs/intro), and infers the schema and data types for you. For an API, start with `dlt init rest_api duckdb`: you [declare the endpoints, pagination, and authentication](https://dlthub.com/docs/dlt-ecosystem/verified-sources/rest_api/basic), and the REST API source does the rest. Build and test on DuckDB locally, then [switch the destination](https://dlthub.com/docs/reference/explainers/how-dlt-works) when you deploy. To [pick a write disposition](https://dlthub.com/docs/general-usage/incremental-loading#how-to-choose-the-right-write-disposition), ask whether your data can change: append records that never change, and merge the ones that do. Keep credentials in `secrets.toml` or environment variables, and [never commit `secrets.toml`](https://dlthub.com/docs/general-usage/credentials/setup#secretstoml-and-configtoml).
|
||||||
|
|
||||||
|
yfinance [fetches market data from Yahoo Finance](https://github.com/ranaroussi/yfinance). Use `Ticker` for one symbol, from its price history to its financial statements, and [`download()` for several symbols](https://ranaroussi.github.io/yfinance/) at once.
|
||||||
|
|
||||||
|
awswrangler is the AWS SDK for pandas. It [connects DataFrames to AWS data and analytics services](https://aws-sdk-pandas.readthedocs.io/en/stable/about.html) like Athena, Glue, Redshift, and S3. It [leaves credentials to boto3 sessions](https://aws-sdk-pandas.readthedocs.io/en/stable/tutorials/002%20-%20Sessions.html): pass your own as `boto3_session`, or it uses the default session. Write S3 data with `wr.s3.to_parquet(..., dataset=True, database=..., table=...)`, and the dataset [goes into the Glue Catalog](https://aws-sdk-pandas.readthedocs.io/en/stable/stubs/awswrangler.s3.to_parquet.html), where Athena can query it.
|
||||||
|
|
||||||
|
Pathway is a [Python ETL framework for stream processing](https://github.com/pathwaycom/pathway), and a Rust engine runs your Python code. To switch between batch and streaming, you [change only the data sources](https://pathway.com/developers/user-guide/introduction/batch-processing/), and the rest of the pipeline stays the same. Streaming is the standard way to run it. Data flows in once you call `pw.run()` and nothing after that call runs, so [read the results through output connectors](https://pathway.com/developers/user-guide/introduction/streaming-and-static-modes/). Pathway is under the [Business Source License](https://pathway.com/developers/user-guide/introduction/licensing-guide/), not an open-source one: production use is free within its limits, and the code converts to Apache after 4 years.
|
||||||
|
|
||||||
|
AKShare is a [Python library for financial data](https://akshare.akfamily.xyz/introduction.html) on stocks, futures, options, funds, bonds, and more, collected from public websites as you call it. Its [stock data](https://akshare.akfamily.xyz/data/stock/stock.html) centers on China, from A-shares and B-shares to the STAR Market, plus Hong Kong and US stocks. Each dataset is one function, like `ak.stock_zh_a_hist()` for A-share daily prices. Its interfaces break when those websites change, so [upgrade AKShare before you use it](https://akshare.akfamily.xyz/installation.html).
|
||||||
|
|
||||||
|
EdgarTools [makes SEC filings easy to access and analyze](https://edgartools.readthedocs.io/en/latest/). Start from a `Company` or a `Filing`, and `.obj()` gives you [a typed object for that form](https://github.com/dgunning/edgartools#how-it-works), with its data as pandas DataFrames. EDGAR requires an email with every request, so set your identity first, with `set_identity()` or the [`EDGAR_IDENTITY` variable](https://edgartools.readthedocs.io/en/latest/configuration/). For bulk work, [turn on local storage](https://edgartools.readthedocs.io/en/latest/guides/local-storage/), which cuts down requests and respects the SEC's rate limits.
|
||||||
|
|
||||||
|
OpenBB's Open Data Platform [integrates proprietary, licensed, and public data sources](https://github.com/OpenBB-finance/OpenBB) and serves them to Python, a REST API, Excel, and MCP servers. Install it in a [new environment](https://docs.openbb.co/odp/python/installation), not your system Python. Pass `provider` to each query: without it, OpenBB [picks the first available provider in alphabetical order](https://docs.openbb.co/odp/python/quickstart). Most providers [need your own API key](https://docs.openbb.co/odp/python/settings/user_settings/api_keys).
|
||||||
|
|
||||||
|
Read the data's terms before you build a product on it, and check its numbers before you trade on them. yfinance's docs say [Yahoo's API is for personal use only](https://ranaroussi.github.io/yfinance/), AKShare says its data is [only for academic research](https://github.com/akfamily/akshare#statement), and OpenBB says its data is [not necessarily accurate](https://github.com/OpenBB-finance/OpenBB#4-disclaimer).
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
Validate API input and config with Pydantic, the Python data validation library built on type hints. Use Pandera for dataframes, jsonschema for JSON Schema.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- API input, forms, and config: Pydantic
|
||||||
|
- Data checked against a JSON Schema document: jsonschema
|
||||||
|
- pandas, polars, or PySpark dataframes: Pandera
|
||||||
|
|
||||||
|
Pydantic builds the schema from your [type hints](https://pydantic.dev/docs/validation/latest/get-started/why/) and guarantees the types of the [output, not the input](https://pydantic.dev/docs/validation/latest/concepts/models/): by default, a numeric string passed to an int field comes out as an int. Where a wrong type should raise an error instead, turn on [strict mode](https://pydantic.dev/docs/validation/latest/concepts/strict_mode/) per field or per model. [Validate incoming JSON directly](https://pydantic.dev/docs/validation/latest/concepts/performance/) instead of parsing it into a dict first. To load config from environment variables, use [Pydantic Settings](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/).
|
||||||
|
|
||||||
|
jsonschema is an [implementation of the JSON Schema specification](https://python-jsonschema.readthedocs.io/en/stable/), so one schema can [work across different systems and platforms](https://json-schema.org/overview/what-is-jsonschema). When you validate many instances against one schema, [create a validator for your schema's draft once and call its `validate` method](https://python-jsonschema.readthedocs.io/en/stable/validate/). Use [`iter_errors()`](https://python-jsonschema.readthedocs.io/en/stable/errors/) to report every error, not only the first. To enforce `format` keywords such as dates or emails, [hook a format checker](https://python-jsonschema.readthedocs.io/en/stable/validate/#validating-formats) into the validator.
|
||||||
|
|
||||||
|
Pandera validates [dataframe-like objects](https://pandera.readthedocs.io/en/stable/): define a schema once and use it on pandas, polars, PySpark, and other dataframe libraries. Write the schema as a DataFrameModel class, [much like a Pydantic model](https://pandera.readthedocs.io/en/stable/dataframe_models.html), and add the `check_types()` decorator to validate at run time. To check an existing pipeline, put [`check_input()` and `check_output()`](https://pandera.readthedocs.io/en/stable/decorators.html) on its functions. To see every failure in one run instead of only the first, validate with [`lazy=True`](https://pandera.readthedocs.io/en/stable/lazy_validation.html).
|
||||||
|
|
||||||
|
Pick by the shape of your data. Running a dataframe through a Pydantic model row by row [might not scale](https://pandera.readthedocs.io/en/stable/pydantic_integration.html) to larger datasets, so use Pandera there; a DataFrameModel can still be a field in a Pydantic model. Pydantic can [generate a JSON Schema](https://pydantic.dev/docs/validation/latest/concepts/json_schema/) from any model for tools that read the format. jsonschema works the other way: it validates data against a JSON Schema document you already have.
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
Pick your Python data visualization library by where the chart goes. Papers take Matplotlib figures, web pages take plotly charts, and data apps use Streamlit.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Figures for papers and reports: Matplotlib
|
||||||
|
- Interactive charts on a web page: plotly
|
||||||
|
- Data dashboards and apps: Streamlit
|
||||||
|
- Statistical charts from a dataframe: seaborn
|
||||||
|
- Charts declared from dataframe columns: Vega-Altair
|
||||||
|
- Interactive plots that call back into Python: Bokeh
|
||||||
|
- Maps in any projection: Cartopy
|
||||||
|
- Graph diagrams laid out by Graphviz: PyGraphviz
|
||||||
|
- Knowledge graph of a codebase: graphify
|
||||||
|
- Demos for machine learning models: Gradio
|
||||||
|
|
||||||
|
Matplotlib has two interfaces. For complicated plots and code you reuse, its docs [suggest the explicit, object-oriented one](https://matplotlib.org/stable/users/explain/quick_start.html#the-explicit-and-the-implicit-interfaces): create the figure with `fig, ax = plt.subplots()`, then call methods on `ax`. The implicit pyplot style is fine for quick interactive work. When you make the same plot for many datasets, write a function that takes the `ax` to draw on.
|
||||||
|
|
||||||
|
plotly draws interactive charts in the browser with plotly.js. Start with [Plotly Express](https://plotly.com/python/plotly-express/), which its docs call the recommended starting point for most common figures: pass a DataFrame and column names, and one call builds the figure. Drop to `go.Figure` for figures Plotly Express can't make or makes awkward, like [subplots of different types or dual-axis plots](https://plotly.com/python/graph-objects/#When-to-use-Graph-Objects-vs-Plotly-Express). To share a chart, [`write_html`](https://plotly.com/python/interactive-html-export/) saves it as an HTML file that stays interactive in any browser.
|
||||||
|
|
||||||
|
seaborn builds statistical graphics on Matplotlib and works on whole datasets. Its docs [recommend the figure-level functions](https://seaborn.pydata.org/tutorial/function_overview.html#relative-merits-of-figure-level-functions), like `relplot()`, for most plots. For one figure that combines different kinds of plots, set it up in Matplotlib and fill it in with axes-level functions. Keep your data in [long form](https://seaborn.pydata.org/tutorial/data_structure.html), one column per variable and one row per observation, as most of seaborn's examples do.
|
||||||
|
|
||||||
|
Vega-Altair is declarative: you [link data columns to encoding channels](https://altair-viz.github.io/getting_started/overview.html) like the x-axis, y-axis, and color, and it handles the rest on top of Vega-Lite. Its docs say the API is [more limited than Matplotlib's or Bokeh's](https://altair-viz.github.io/getting_started/project_philosophy.html), a trade they make to keep exploring data simple. A chart carries its data inside the spec, so for a large dataset, [enable the VegaFusion data transformer](https://altair-viz.github.io/user_guide/large_datasets.html#vegafusion-data-transformer) or pass the data by URL.
|
||||||
|
|
||||||
|
Bokeh builds interactive plots for the browser without any JavaScript from you. Start with [`bokeh.plotting`, its primary interface](https://docs.bokeh.org/en/latest/docs/user_guide/intro.html#the-bokeh-plotting-interface), and `output_file()` for a standalone HTML file or `output_notebook()` for Jupyter. What sets it apart is the [Bokeh server](https://docs.bokeh.org/en/latest/docs/user_guide/server/server_introduction.html), which keeps data in sync between Python and the browser. Widgets can then run Python callbacks, and plots can stream data. Write the app as a script and [serve it with `bokeh serve`](https://docs.bokeh.org/en/latest/docs/user_guide/server/app.html#building-applications).
|
||||||
|
|
||||||
|
Cartopy draws maps on Matplotlib and [suits data over large areas](https://cartopy.readthedocs.io/stable/), where Cartesian math breaks down at the poles and the dateline. Set the map's projection on the axes, and [always pass `transform`](https://cartopy.readthedocs.io/stable/tutorials/understanding_transform.html) to say which coordinate system your data is in.
|
||||||
|
|
||||||
|
PyGraphviz is a Python interface to Graphviz. Build a graph with `AGraph` or read a DOT file into one, then [lay it out and draw it](https://pygraphviz.github.io/documentation/stable/tutorial.html#layout-and-drawing) with one of Graphviz's layout programs.
|
||||||
|
|
||||||
|
graphify maps a project's code, docs, PDFs, and images into a [knowledge graph your coding agent can query](https://github.com/Graphify-Labs/graphify). It draws the graph as a `graph.html` you can click through in a browser. It parses code locally, so code never leaves your machine; docs and media go through your agent's model. Install the `graphifyy` package, run `graphify install`, then type `/graphify .` in your agent.
|
||||||
|
|
||||||
|
Streamlit turns a Python script into a data app, and [reruns the whole script from top to bottom](https://docs.streamlit.io/get-started/fundamentals/main-concepts) every time something on screen changes. Start it with `streamlit run`. To skip repeated work on those reruns, [cache](https://docs.streamlit.io/get-started/fundamentals/advanced-concepts) data with `st.cache_data`, and shared resources like ML models or database connections with `st.cache_resource`. Keep per-user values in Session State.
|
||||||
|
|
||||||
|
Gradio wraps a Python function, often a machine learning model, in a web UI. [Use `gr.Interface`](https://gradio.app/guides/quickstart) for a demo with inputs and outputs, `gr.ChatInterface` for a chatbot, and `gr.Blocks` for custom layouts and data flows. `launch(share=True)` gives you a public link, and [Hugging Face Spaces](https://gradio.app/guides/sharing-your-app#hosting-on-hf-spaces) hosts the app for good. Anyone with the link can call your function, so [put a login in front of it](https://gradio.app/guides/sharing-your-app#password-protected-app) or keep sensitive data out. Its security docs also recommend that you [set `max_file_size` and keep `allowed_paths` as small as possible](https://gradio.app/guides/file-access#best-practices).
|
||||||
|
|
||||||
|
seaborn and Cartopy draw on Matplotlib Axes, so to [customize what they draw](https://matplotlib.org/stable/users/explain/figure/api_interfaces.html), use Matplotlib's explicit Axes interface. Streamlit shows [Matplotlib, plotly, and Vega-Altair figures](https://docs.streamlit.io/develop/api-reference/charts), and Gradio's [`gr.Plot`](https://gradio.app/docs/gradio/plot) takes those plus Bokeh, so your plotting code carries over into an app.
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
PostgreSQL gets Psycopg, MySQL mysqlclient, and SQLite the built-in sqlite3. Elsewhere, use the Python database driver from your database's maker.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- PostgreSQL, from sync or async code: Psycopg
|
||||||
|
- MySQL or MariaDB: mysqlclient, or PyMySQL for pure Python under the MIT license
|
||||||
|
- SQLite in your app: sqlite3
|
||||||
|
- PostgreSQL from asyncio code, when query speed comes first: asyncpg
|
||||||
|
- Loading JSON or CSV into SQLite and reshaping its tables, from the shell or Python: sqlite-utils
|
||||||
|
- ClickHouse: ClickHouse Connect, or clickhouse-driver for the native TCP protocol
|
||||||
|
- Any other database with an ODBC driver: pyodbc
|
||||||
|
- Oracle Database: python-oracledb
|
||||||
|
- SQL Server or Azure SQL, with no driver manager to install: mssql-python
|
||||||
|
- Redis: redis-py
|
||||||
|
- MongoDB: PyMongo, or Django MongoDB Backend in a Django project
|
||||||
|
- Apache Cassandra: cassandra-driver
|
||||||
|
|
||||||
|
Psycopg keeps the [DB-API interface](https://www.psycopg.org/psycopg3/docs/) of the older Psycopg and adds asyncio support, so sync and async code share one driver. Open a connection [in a `with` block](https://www.psycopg.org/psycopg3/docs/basic/usage.html#connection-context): it commits when the block ends, rolls back if an exception is raised, and closes the connection either way. When several threads need connections, take them from a [`ConnectionPool`](https://www.psycopg.org/psycopg3/docs/advanced/pool.html#basic-connection-pool-usage), or an `AsyncConnectionPool` in async code. Django [recommends Psycopg](https://docs.djangoproject.com/en/stable/ref/databases/#postgresql-notes) for PostgreSQL.
|
||||||
|
|
||||||
|
asyncpg is built for asyncio and [speaks PostgreSQL's protocol natively](https://github.com/MagicStack/asyncpg) instead of hiding it behind the DB-API, so its API is its own, down to `$1` placeholders. In a server, [use its connection pool](https://magicstack.github.io/asyncpg/current/usage.html#connection-pools): take a connection per request with `async with pool.acquire()`, and wrap writes in `async with connection.transaction()`. Outside a transaction, each statement commits right away.
|
||||||
|
|
||||||
|
mysqlclient is a native driver that builds against the MySQL client library, and it's [Django's recommended choice](https://docs.djangoproject.com/en/stable/ref/databases/#mysql-db-api-drivers) for MySQL. PyMySQL is [pure Python](https://github.com/PyMySQL/PyMySQL), so it installs without that library. The licenses differ, too: mysqlclient is [GPL](https://github.com/PyMySQL/mysqlclient/blob/main/LICENSE), and PyMySQL is [MIT](https://github.com/PyMySQL/PyMySQL/blob/main/LICENSE).
|
||||||
|
|
||||||
|
sqlite3 ships with Python, and its docs suggest SQLite for an app's internal storage, or for [a prototype you later port](https://docs.python.org/3/library/sqlite3.html) to a larger database like PostgreSQL. Use the connection [as a context manager](https://docs.python.org/3/library/sqlite3.html#sqlite3-connection-context-manager) to commit or roll back a transaction; it doesn't close the connection, so close it yourself. sqlite-utils is [not a full ORM](https://sqlite-utils.datasette.io/en/stable/) but a set of helpers for creating a SQLite database and filling it with data, from Python or its command line. [Pipe JSON or CSV into it](https://github.com/simonw/sqlite-utils), and it creates the table for you. It also runs schema changes that SQLite's `ALTER TABLE` can't, like changing a column's type.
|
||||||
|
|
||||||
|
ClickHouse Connect is the Python driver that [ClickHouse's own docs](https://clickhouse.com/docs/integrations/language-clients/python/index) cover. It has a sync and an async client over the HTTP interface, which works through load balancers and proxies. clickhouse-driver talks ClickHouse's [native TCP protocol](https://clickhouse-driver.readthedocs.io/en/latest/) instead. To insert rows fast with it, [pass them separately](https://clickhouse-driver.readthedocs.io/en/latest/quickstart.html#inserting-data) and end the statement with `VALUES`.
|
||||||
|
|
||||||
|
pyodbc connects to [any database with an ODBC driver](https://github.com/mkleehammer/pyodbc). On macOS and Linux, install an ODBC driver manager like unixODBC first; Windows has one built in. python-oracledb is Oracle's own driver and the [successor to cx_Oracle](https://python-oracledb.readthedocs.io/en/latest/user_guide/introduction.html). Its default Thin mode [connects without Oracle Client libraries](https://python-oracledb.readthedocs.io/en/latest/user_guide/initialization.html), which is enough for most apps. mssql-python is [Microsoft's own driver](https://learn.microsoft.com/en-us/sql/connect/python/mssql-python/migrate-from-pyodbc) for SQL Server and Azure SQL. It [connects without an external driver manager](https://learn.microsoft.com/en-us/sql/connect/python/mssql-python/python-sql-driver-mssql-python) and [pools connections by default](https://github.com/microsoft/mssql-python#connection-pooling).
|
||||||
|
|
||||||
|
redis-py is [the Python client for Redis](https://redis.io/docs/latest/develop/clients/redis-py/), with asyncio support. PyMongo is [the recommended way](https://www.mongodb.com/docs/languages/python/pymongo-driver/current/) to work with MongoDB from Python: use `MongoClient` in sync code and [`AsyncMongoClient`](https://www.mongodb.com/docs/languages/python/pymongo-driver/current/connect/mongoclient/) in async code. Django MongoDB Backend is [a Django database backend](https://github.com/mongodb/django-mongodb-backend) that uses PyMongo, so your Django models live in MongoDB. Joins [don't perform well on large tables](https://django-mongodb-backend.readthedocs.io/en/latest/faq/#performance) there, so model related data as embedded models. With cassandra-driver, [use prepared statements](https://docs.datastax.com/en/developer/python-driver/latest/getting_started/#prepared-statement) for queries you run often, so Cassandra doesn't parse them again each time.
|
||||||
|
|
||||||
|
Whatever the database, create the client or pool once per process and share it. ClickHouse Connect's docs say to [create clients once at startup](https://clickhouse.com/docs/integrations/language-clients/python/driver-api#client-lifecycle-and-best-practices), and a PyMongo `MongoClient` or a redis-py `Redis` object already holds a pool. If your server forks worker processes, [create the client or pool after the fork](https://www.psycopg.org/psycopg3/docs/advanced/async.html#concurrent-operations). Pass values as query parameters. Building SQL strings yourself [opens the door to SQL injection](https://docs.python.org/3/library/sqlite3.html#sqlite3-placeholders).
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
No server to run: a Python database library can live inside your process. Pick DuckDB when you run analytical SQL, LanceDB when you search vectors.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Analytical SQL over DataFrames and Parquet files: DuckDB
|
||||||
|
- Vector search with the raw data stored beside its embeddings: LanceDB
|
||||||
|
- ClickHouse SQL in a notebook, the same queries your ClickHouse server runs: chDB
|
||||||
|
- A RAG prototype that embeds your text for you: Chroma
|
||||||
|
- Vector search, full-text search, and filters in one query: Zvec
|
||||||
|
- Tables whose columns run AI models on every insert: Pixeltable
|
||||||
|
- Python dicts in a JSON file, for a small app with one process: TinyDB
|
||||||
|
|
||||||
|
DuckDB is built for [analytical queries](https://duckdb.org/why_duckdb), and it runs inside your Python process with no database server to install. `duckdb.sql()` runs on an in-memory database, while [`duckdb.connect()` with a file name](https://duckdb.org/docs/current/clients/python/overview#persistent-storage) keeps your tables on disk. It [queries pandas and Polars DataFrames and Arrow tables directly](https://duckdb.org/docs/current/clients/python/overview#dataframes), and Parquet files by their file name. In a package others import, [create your own connection objects](https://duckdb.org/docs/current/clients/python/overview#connection-object-and-module) instead of calling the module's functions, which share one global database.
|
||||||
|
|
||||||
|
chDB is an [in-process SQL engine powered by ClickHouse](https://clickhouse.com/docs/chdb), for ClickHouse SQL without a ClickHouse server. It speaks [the full ClickHouse SQL dialect](https://clickhouse.com/resources/engineering/what-is-chdb), so a query you write in a notebook runs unchanged on a ClickHouse server later. `chdb.query()` is stateless. For tables that last across queries, open a `Session`, and [give it a directory name](https://clickhouse.com/docs/chdb/getting-started#creating-a-table-from-json-file) to keep them on disk.
|
||||||
|
|
||||||
|
Chroma's embedded mode is for [prototyping and experimentation](https://docs.trychroma.com/reference/architecture/overview). It gets a RAG prototype running fast, since it [embeds and indexes your documents for you](https://docs.trychroma.com/docs/overview/getting-started) with a default model that runs on your machine. Save data to a directory with [`PersistentClient`](https://docs.trychroma.com/docs/run-chroma/clients#persistent-client). For production, its docs [prefer a Chroma server](https://docs.trychroma.com/reference/python/client#persistentclient) that your app connects to as a client.
|
||||||
|
|
||||||
|
LanceDB is an [embedded retrieval library](https://docs.lancedb.com/) that runs in your process: [point it at a local directory](https://docs.lancedb.com/quickstart#connect-via-local-directory-path), or at an object storage URI like `s3://`. It [stores the raw data, metadata, and embeddings together](https://docs.lancedb.com/faq/faq-oss#what-makes-lancedb-different), and its indexes live on disk.
|
||||||
|
|
||||||
|
Zvec is a vector database that [runs entirely in-process](https://zvec.org/en/docs/db/), from notebooks and servers to edge devices. [Define a schema](https://zvec.org/en/docs/db/quickstart/) with scalar fields and vectors, then create a collection from it. One query can [combine vector similarity, full-text search, and filters](https://github.com/alibaba/zvec#user-content--features).
|
||||||
|
|
||||||
|
Pixeltable is more than a vector store: it's [the database, orchestration, and serving](https://docs.pixeltable.com/overview/pixeltable) in one Python file. [Model inference can go in a computed column](https://docs.pixeltable.com/tutorials/computed-columns), which [runs on insert and on update](https://docs.pixeltable.com/overview/how-it-works). Each new row gets its model outputs without a pipeline you rerun. Put an embedding index on a column, and [each insert keeps it current](https://docs.pixeltable.com/howto/coming-from), with no separate vector database.
|
||||||
|
|
||||||
|
TinyDB is pure Python with no dependencies, and [`TinyDB('db.json')`](https://tinydb.readthedocs.io/en/latest/getting-started.html) gives you a database that stores Python dicts in that JSON file. Its docs call it [the wrong database](https://tinydb.readthedocs.io/en/latest/intro.html#why-not-use-tinydb) when you need access from several processes or threads, indexes, ACID guarantees, or high performance.
|
||||||
|
|
||||||
|
Most of these databases expect one process to write at a time. DuckDB lets [one process read and write](https://duckdb.org/docs/current/connect/concurrency#single-process), or several processes only read. A chDB data directory [opens in one process at a time](https://clickhouse.com/docs/chdb/getting-started). Zvec shares a collection across processes [in read-only mode](https://zvec.org/en/docs/db/collections/open/), and Pixeltable keeps [one writer process](https://docs.pixeltable.com/howto/deployment/operations) even with several API workers.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
Most code should handle time zones with zoneinfo and pick python-dateutil, the Python date library for parsing date strings and adding months.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Time zones by IANA name, like Europe/Paris: zoneinfo
|
||||||
|
- Parsing date strings, adding months, and recurring dates: python-dateutil
|
||||||
|
- Dates as people write them, like "3 days ago", in many languages: dateparser
|
||||||
|
- An easier datetime API whose objects are still datetimes: pendulum
|
||||||
|
- Exact and local times as separate types, with DST-safe math: whenever
|
||||||
|
|
||||||
|
zoneinfo brings the IANA time zone database to Python, and the datetime docs say [its usage is recommended](https://docs.python.org/3/library/datetime.html#tzinfo-objects). Attach a ZoneInfo to a datetime [through the constructor, `replace()`, or `astimezone()`](https://docs.python.org/3/library/zoneinfo.html#using-zoneinfo). Some systems, Windows among them, have no IANA database, so if your code runs across platforms, [declare a dependency on tzdata](https://docs.python.org/3/library/zoneinfo.html#data-sources).
|
||||||
|
|
||||||
|
python-dateutil adds [extensions to the standard datetime module](https://dateutil.readthedocs.io/en/stable/); install it as `python-dateutil` and import it as `dateutil`. Its parser [reads most known formats](https://dateutil.readthedocs.io/en/stable/parser.html) and returns a datetime even for an ambiguous date. For input like 01/05/09, set `dayfirst` or `yearfirst` to match your data. For calendar math, `relativedelta` takes [plural arguments that add and singular ones that replace](https://dateutil.readthedocs.io/en/stable/relativedelta.html): `months=+1` moves a month ahead, and `day=1` jumps to the first. `rrule` builds [recurring dates from iCalendar rules](https://dateutil.readthedocs.io/en/stable/rrule.html).
|
||||||
|
|
||||||
|
dateparser reads relative dates like "two weeks ago" and absolute ones in more than 200 language locales. Its docs say it [stands out](https://dateparser.readthedocs.io/en/latest/#common-use-cases) for scraped pages, logs, and other data from mixed sources, and for letting users type dates in their own words. Call `dateparser.parse()`, and [pass `languages` when you know them](https://dateparser.readthedocs.io/en/latest/#how-to-use), so it skips language detection. When you parse many dates from one source, [use `DateDataParser`](https://dateparser.readthedocs.io/en/latest/usage.html), which remembers the languages it has found.
|
||||||
|
|
||||||
|
pendulum's classes are [drop-in replacements for the native ones](https://pendulum.eustace.io/docs/#introduction), since they inherit from datetime. Every instance is time zone aware and in UTC by default. Its docs call aware datetimes [the preferred and recommended way](https://pendulum.eustace.io/docs/#instantiation) to use it. For tests, install `pendulum[test]` and [travel in time](https://pendulum.eustace.io/docs/#testing).
|
||||||
|
|
||||||
|
whenever puts exact time and local time in [separate types](https://whenever.readthedocs.io/en/latest/guide/choosing-a-type.html): an instant when only the moment matters, a zoned datetime when the local time matters too. Mixing up naive and aware [becomes a type error](https://whenever.readthedocs.io/en/latest/), and DST is handled in all arithmetic. A standard datetime [does no time zone adjustment](https://docs.python.org/3/library/datetime.html#datetime-objects) when you add a timedelta to it. In production, [turn whenever's DST warnings into errors](https://whenever.readthedocs.io/en/latest/faq.html#why-warnings-instead-of-errors) with Python's standard warnings filter.
|
||||||
|
|
||||||
|
Decide whether you extend datetime or replace it. zoneinfo, python-dateutil, and dateparser all use standard datetime objects, so they work together. pendulum's objects are datetimes too, but code that checks the exact type, like sqlite3 and some database drivers, [needs an adapter registered](https://pendulum.eustace.io/docs/#limitations). whenever [doesn't subclass datetime at all](https://whenever.readthedocs.io/en/latest/faq.html#why-no-drop-in-replacement-for-datetime), so [convert to and from standard datetimes](https://whenever.readthedocs.io/en/latest/guide/stdlib-convert.html) where other code needs one.
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
Before you add another print call, try a Python debugging tool. Step through your code in ipdb, and when it's slow, find where the time goes with py-spy.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Stepping through code with pdb's commands: ipdb
|
||||||
|
- Profiling without changing code, even in production: py-spy
|
||||||
|
- A full-screen debugger in the terminal: PuDB
|
||||||
|
- Tracing calls through a big application: Hunter
|
||||||
|
- Memory use and leaks, on Linux and macOS: Memray
|
||||||
|
- A call tree that includes time spent waiting on I/O: pyinstrument
|
||||||
|
- CPU, GPU, and memory profiles line by line: Scalene
|
||||||
|
- Debug panels in a Django or Flask app: Django Debug Toolbar or Flask-DebugToolbar
|
||||||
|
- Print debugging that shows each expression: IceCream
|
||||||
|
|
||||||
|
ipdb gives you the IPython debugger with [the same interface as pdb](https://github.com/gotcha/ipdb). It adds tab completion, syntax highlighting, and better tracebacks. Call `ipdb.set_trace()` where you want to stop and look around.
|
||||||
|
|
||||||
|
PuDB is a full-screen debugger that runs in your terminal, with the source, the stack, breakpoints, and variables [all visible at once](https://documen.tician.de/pudb/). Call `from pudb import set_trace; set_trace()` where you want to stop, or [run a whole script](https://documen.tician.de/pudb/starting.html) under it with `python -m pudb my-script.py`.
|
||||||
|
|
||||||
|
Hunter traces what your code does, to help you [understand and debug big applications](https://github.com/ionelmc/python-hunter). Its main selling point is filtering the events you see. Start it from code with `hunter.trace()`, from the `PYTHONHUNTER` environment variable, or with the `hunter-trace` CLI, which [attaches to a running process](https://python-hunter.readthedocs.io/en/latest/introduction.html#activation). To see only your own code, [set `stdlib=False`](https://python-hunter.readthedocs.io/en/latest/cookbook.html#typical).
|
||||||
|
|
||||||
|
py-spy shows where your program spends its time [without restarting it or changing its code](https://github.com/benfred/py-spy). It runs outside your program's process, so its docs call it safe to use on production code. `py-spy record -o profile.svg --pid 12345` writes a flame graph of a running process, or pass `-- python myprogram.py` in place of the PID to start one. When a program hangs, `py-spy dump` prints its current call stack.
|
||||||
|
|
||||||
|
pyinstrument records [wall-clock time](https://pyinstrument.readthedocs.io/en/latest/how-it-works.html#wall-clock-time-not-cpu-time), so the time your program spends downloading data, reading files, and talking to databases shows up in its call tree. It samples the call stack instead of tracing every call, which [keeps its overhead low](https://pyinstrument.readthedocs.io/en/latest/how-it-works.html#statistical-profiling-not-tracing). Type `pyinstrument script.py` instead of `python script.py`, or [wrap the code you want to profile](https://pyinstrument.readthedocs.io/en/latest/guide.html#profile-a-specific-chunk-of-code) in a `with pyinstrument.profile():` block.
|
||||||
|
|
||||||
|
Scalene profiles CPU, GPU, and memory [line by line](https://github.com/plasma-umass/scalene). It separates the time spent in Python from the time spent in native code, so you can focus on the code you can actually improve. It also points to the lines responsible for memory growth and likely leaks.
|
||||||
|
|
||||||
|
Memray tracks memory allocations [in Python code, native extension modules, and the interpreter itself](https://bloomberg.github.io/memray/overview.html). It traces every function call rather than sampling, so the call stacks it reports are accurate. Use it to find what's using memory, where it leaks, and which code allocates the most. Profile [in two steps](https://bloomberg.github.io/memray/getting_started.html): `memray run example.py` saves the allocations to a file, and `memray flamegraph` turns that file into a report. Memray only works on Linux and macOS, so on Windows, profile memory with Scalene.
|
||||||
|
|
||||||
|
Django Debug Toolbar adds panels with debug information about the current request and response. [Set it up](https://django-debug-toolbar.readthedocs.io/en/latest/installation.html) by adding its app, URLs, and middleware. The toolbar shows only for the IP addresses in `INTERNAL_IPS`, so add `"127.0.0.1"` there. Its docs warn that it [isn't hardened for production](https://django-debug-toolbar.readthedocs.io/en/latest/configuration.html#show-toolbar-callback) or public servers. Flask-DebugToolbar is [a port of it](https://github.com/pallets-eco/flask-debugtoolbar) for Flask: pass your app to `DebugToolbarExtension(app)`, and the toolbar [appears in HTML responses when debug mode is on](https://flask-debugtoolbar.readthedocs.io/en/latest/#usage).
|
||||||
|
|
||||||
|
IceCream's `ic()` is [like `print()`, but better](https://github.com/gruns/icecream): `ic(foo(123))` prints both the expression and its value: `ic| foo(123): 456`. With no arguments, it prints the file, line number, and function it's called from. It returns its arguments, so you can wrap it around code that's already there. When you're done, `ic.disable()` turns off all its output.
|
||||||
|
|
||||||
|
ipdb and PuDB both keep the interface of pdb, the debugger that ships with Python. You don't have to import either one in your code. Call the built-in `breakpoint()` instead, and set the `PYTHONBREAKPOINT` environment variable to [the function it should run](https://docs.python.org/3/using/cmdline.html#envvar-PYTHONBREAKPOINT), like `ipdb.set_trace` or `pudb.set_trace`. Left unset, `breakpoint()` starts pdb.
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
Three Python deep learning frameworks, three strengths: PyTorch for research and new architectures, Keras for a high-level API, JAX for compiled code on TPUs.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Research and new model architectures: PyTorch
|
||||||
|
- High-level model building on JAX or PyTorch: Keras
|
||||||
|
- NumPy-style array code compiled for GPUs and TPUs: JAX
|
||||||
|
- PyTorch training without the loop boilerplate: PyTorch Lightning
|
||||||
|
- Reinforcement learning environments: Gymnasium
|
||||||
|
- Reinforcement learning algorithms: Stable-Baselines3
|
||||||
|
|
||||||
|
PyTorch models are `nn.Module` subclasses, and autograd [builds the computational graph as your code runs](https://docs.pytorch.org/docs/stable/user_guide/pytorch_main_components.html). Install it with [the command from its selector](https://pytorch.org/get-started/locally/), which matches your OS and GPU. To keep a model, [save its `state_dict`](https://docs.pytorch.org/tutorials/beginner/saving_loading_models.html#save-load-state-dict-recommended) instead of pickling the whole module. Wrap your model in [`torch.compile`](https://docs.pytorch.org/tutorials/intermediate/torch_compile_tutorial.html) to speed it up with minimal code changes. To train on more than one GPU, use [DistributedDataParallel](https://docs.pytorch.org/docs/stable/notes/cuda.html#use-nn-parallel-distributeddataparallel-instead-of-multiprocessing-or-nn-dataparallel).
|
||||||
|
|
||||||
|
Keras is a [multi-framework API](https://keras.io/getting_started/about/): a Keras model can run as a PyTorch Module or as a JAX function. Install a backend next to it and [set `KERAS_BACKEND`](https://keras.io/getting_started/#configuring-your-backend) before you import Keras. Build simple models as a Sequential stack of layers and anything more complex with the functional API. Save the whole model to [a `.keras` file](https://keras.io/getting_started/faq/#what-are-my-options-for-saving-models) with `model.save()` instead of pickling it. The file [reloads with any backend](https://keras.io/keras_3/).
|
||||||
|
|
||||||
|
JAX does [accelerator-oriented array computation](https://docs.jax.dev/en/latest/) with a NumPy-style API and composable transformations: `jax.grad` for derivatives, `jax.jit` for compilation, and `jax.vmap` for batching. The transformations only work on [functionally pure functions](https://docs.jax.dev/en/latest/notebooks/Common_Gotchas_in_JAX.html#pure-functions), so pass all data in as arguments and return every result. JAX itself stays narrow: to train neural networks, use the [JAX AI Stack](https://docs.jaxstack.ai/en/latest/getting_started.html), with Flax NNX for models and Optax for optimizers. On NVIDIA GPUs, the JAX team strongly recommends [installing CUDA and cuDNN from pip wheels](https://docs.jax.dev/en/latest/installation.html#pip-installation-nvidia-gpu-cuda-installed-via-pip-easier).
|
||||||
|
|
||||||
|
PyTorch Lightning [organizes PyTorch code to remove boilerplate](https://lightning.ai/docs/pytorch/stable/home/introduction): you write the model logic in a LightningModule, and the Trainer handles devices, precision, and distributed training. Install it as [the `lightning` package](https://lightning.ai/docs/pytorch/stable/home/installation). Its [style guide](https://lightning.ai/docs/pytorch/stable/reference/starter/style_guide) recommends keeping each LightningModule self-contained, the model separate from the system that trains it, and data loading in a LightningDataModule. To keep your own training loop, use [Lightning Fabric](https://lightning.ai/docs/fabric/stable), which scales a plain PyTorch script after you change a few lines.
|
||||||
|
|
||||||
|
Gymnasium is [an API standard for reinforcement learning](https://gymnasium.farama.org/), with a collection of reference environments. It's the maintained fork of OpenAI's Gym, and many older tutorials still use Gym's old API, so follow its [migration guide](https://gymnasium.farama.org/introduction/migration_guide/) when you port one. [Register your own environment](https://gymnasium.farama.org/introduction/create_custom_env/#registering-and-making-the-environment) so `gymnasium.make()` creates it like a built-in one, and run [`check_env`](https://gymnasium.farama.org/introduction/create_custom_env/#check-environment-validity) on it to catch common issues.
|
||||||
|
|
||||||
|
Stable-Baselines3 is a set of [reliable implementations of reinforcement learning algorithms in PyTorch](https://stable-baselines3.readthedocs.io/en/master/), and it trains on any environment that [follows the Gymnasium interface](https://stable-baselines3.readthedocs.io/en/master/guide/custom_env.html). It [assumes you know some reinforcement learning](https://github.com/DLR-RM/stable-baselines3). Its tips page recommends [starting from the RL Zoo's tuned hyperparameters](https://stable-baselines3.readthedocs.io/en/master/guide/rl_tips.html#general-advice-when-using-reinforcement-learning) and normalizing the agent's input. Evaluate the agent on [a separate test environment](https://stable-baselines3.readthedocs.io/en/master/guide/rl_tips.html#how-to-evaluate-an-rl-algorithm), since training adds exploration noise. [Pick an algorithm](https://stable-baselines3.readthedocs.io/en/master/guide/rl_tips.html#which-algorithm-should-i-use) by your action space first: DQN handles only discrete actions, and SAC only continuous ones.
|
||||||
|
|
||||||
|
PyTorch's security policy says [running untrusted models is equivalent to running untrusted code](https://github.com/pytorch/pytorch/blob/main/SECURITY.md), so run untrusted ones in a sandbox. Load checkpoints with [`weights_only=True`](https://docs.pytorch.org/docs/stable/notes/serialization.html#weights-only-security) in `torch.load`, and leave [`safe_mode`](https://keras.io/api/models/model_saving_apis/model_saving_and_loading/) on when Keras loads a model.
|
||||||
@@ -0,0 +1,46 @@
|
|||||||
|
Across a fleet of servers, Ansible applies YAML playbooks over SSH with no agent to install. Clouds publish their own Python DevOps tools, like Boto3 for AWS.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Servers configured from YAML playbooks over SSH: Ansible
|
||||||
|
- AWS, Azure, or Google Cloud from Python code: Boto3, the Azure SDK for Python, or google-cloud-python
|
||||||
|
- AWS from your terminal and shell scripts: AWS CLI
|
||||||
|
- A new cloud instance set up on its first boot: cloud-init
|
||||||
|
- Server configuration written in Python instead of YAML: pyinfra
|
||||||
|
- An agent on every server, run from a central master: Salt
|
||||||
|
- Shell commands on remote servers, run from Python code: Fabric
|
||||||
|
- Python APIs and event handlers on AWS Lambda: Chalice
|
||||||
|
- Process and system stats, or other programs called as functions: psutil, sh
|
||||||
|
- Your app's errors, or Celery workers and tasks, monitored live: Sentry SDK, Flower
|
||||||
|
- Your app's processes kept running on Unix: Supervisor
|
||||||
|
- Encrypted backups, or chaos engineering experiments: BorgBackup, Chaos Toolkit
|
||||||
|
|
||||||
|
Ansible playbooks [declare the state you want each system in](https://docs.ansible.com/projects/ansible/latest/getting_started/introduction.html), written in YAML. Ansible connects over SSH with your existing credentials, so the servers it manages need no extra software. When a system already matches the playbook, Ansible changes nothing. Run a playbook with `ansible-playbook`, and run it [with `--check` first](https://docs.ansible.com/projects/ansible/latest/playbook_guide/playbooks_intro.html#running-playbooks-in-check-mode) to get a report of the changes it would make, without making them.
|
||||||
|
|
||||||
|
Boto3 is the AWS SDK for Python, and it [shares its low-level core with the AWS CLI](https://docs.aws.amazon.com/boto3/latest/guide/quickstart.html). The AWS CLI calls the same AWS APIs from your shell, for exploring a service and writing shell scripts. [Install it from AWS's own installers](https://docs.aws.amazon.com/cli/latest/userguide/cli-chap-welcome.html): the builds in package managers are unofficial.
|
||||||
|
|
||||||
|
The Azure SDK for Python is a set of separate libraries for specific Azure services. Its [management libraries, named `azure-mgmt-*`](https://learn.microsoft.com/en-us/azure/developer/python/sdk/azure-sdk-overview#create-and-manage-azure-resources-with-management-libraries), create and configure resources, while its client libraries work with resources that already exist. On Google Cloud, google-cloud-python holds the [Cloud Client Libraries, the option Google recommends](https://docs.cloud.google.com/apis/docs/client-libraries-explained) for calling its APIs from code.
|
||||||
|
|
||||||
|
cloud-init gives a new cloud instance [its configuration on first boot](https://docs.cloud-init.io/en/latest/), with nothing to install, and every major public cloud supports it. Write that configuration as [cloud-config](https://docs.cloud-init.io/en/latest/explanation/format/cloud-config.html), YAML whose keys describe the state you want, like packages, users, and SSH keys. For more complex configuration, cloud-init [can hand over to a tool like Ansible](https://docs.cloud-init.io/en/latest/explanation/introduction.html).
|
||||||
|
|
||||||
|
pyinfra [turns Python code into shell commands and runs them on your servers](https://github.com/pyinfra-dev/pyinfra): think Ansible, but Python instead of YAML. A deploy is [an `inventory.py` of hosts and a `deploy.py` of operations](https://docs.pyinfra.com/en/latest/getting-started.html), run with `pyinfra inventory.py deploy.py`. Operations declare a state, like a package being installed, and pyinfra changes only what differs. The target hosts need nothing but an SSH server.
|
||||||
|
|
||||||
|
Salt is [a remote execution framework for configuration management and orchestration](https://docs.saltproject.io/salt/user-guide/en/latest/topics/overview.html). A Salt master sends commands to minions: the systems it manages, each running the salt-minion service. salt-ssh reaches systems without that agent. Still, Salt's docs [recommend the standard install](https://docs.saltproject.io/salt/install-guide/en/latest/topics/overview.html#standard-installation-overview) of a master plus minions for most organizations, since the agentless setup lacks some features.
|
||||||
|
|
||||||
|
Fabric is [a library that runs shell commands over SSH](https://www.fabfile.org/) and returns the results as Python objects, built on Invoke and Paramiko. Open a `Connection` to a host and call `run()` on it. To run your code from the shell, [write `@task` functions in a `fabfile.py`](https://docs.fabfile.org/en/latest/getting-started.html#addendum-the-fab-command-line-tool) and call them with `fab`.
|
||||||
|
|
||||||
|
Chalice is [a framework for serverless apps on AWS](https://aws.github.io/chalice/). Flask-style decorators hook your functions up to HTTP routes, schedules, and S3 events. Then [`chalice deploy`](https://aws.github.io/chalice/quickstart.html) provisions what they need on API Gateway and Lambda.
|
||||||
|
|
||||||
|
psutil [reads process and system stats](https://psutil.io/), like CPU, memory, disks, and network, with one API on every platform it supports. Its docs call [parsing the output of `ps` or `top`](https://psutil.io/alternatives/) fragile, since psutil reads the same kernel data directly. sh [calls any program as if it were a function](https://sh.readthedocs.io/en/latest/), on Unix-like systems only.
|
||||||
|
|
||||||
|
The Sentry SDK reports your app's errors and uncaught exceptions to Sentry. [Initialize it in your app's entry point](https://docs.sentry.io/platforms/python/), as early as possible. For Django, FastAPI, or another web framework, follow that framework's guide instead.
|
||||||
|
|
||||||
|
Supervisor [monitors and controls your project's processes](https://supervisord.org/) on Unix-like systems, without replacing init. Run each program [in the foreground, not as a daemon](https://supervisord.org/subprocess.html#nondaemonizing-of-subprocesses), so Supervisor can control it.
|
||||||
|
|
||||||
|
Flower is a web app showing the status of Celery workers and tasks in real time, and it's [Celery's recommended monitor](https://docs.celeryq.dev/en/stable/userguide/monitoring.html#flower-real-time-celery-web-monitor). [Start it with `celery -A <your app> flower`](https://flower.readthedocs.io/en/latest/install.html).
|
||||||
|
|
||||||
|
BorgBackup makes compressed, deduplicated backups [with authenticated encryption](https://www.borgbackup.org/), so your backup server only ever sees ciphertext. [Keep a copy of your key](https://borgbackup.readthedocs.io/en/stable/quickstart.html#repository-encryption) with `borg key export`.
|
||||||
|
|
||||||
|
Chaos Toolkit runs chaos engineering experiments. [Each experiment declares](https://chaostoolkit.org/reference/concepts/) a steady-state hypothesis, which describes what normal looks like for your system. Its method runs actions and probes, and rollbacks can revert those actions.
|
||||||
|
|
||||||
|
Whatever cloud your code talks to, keep access keys out of it. On the cloud's own machines, use the identity the machine already has: [an IAM role on EC2](https://docs.aws.amazon.com/boto3/latest/guide/credentials.html#best-practices-for-configuring-credentials), [a managed identity on Azure](https://learn.microsoft.com/en-us/azure/developer/python/sdk/authentication/overview), [the attached service account on Google Cloud](https://docs.cloud.google.com/docs/authentication/application-default-credentials#attached-sa). On your own machine, sign in with the cloud's command-line tool. The SDK's [credential chain](https://learn.microsoft.com/en-us/azure/developer/python/sdk/authentication/credential-chains) picks up that login, so the same code runs in both places.
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
From a laptop to a cluster, Python distributed computing comes down to the workload: joblib runs loops, Dask scales pandas, Ray scales ML, and PySpark runs SQL.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Parallel for loops on one machine: joblib
|
||||||
|
- Scaling pandas, NumPy, or scikit-learn code: Dask
|
||||||
|
- Training, tuning, and serving ML models, or GPU jobs: Ray
|
||||||
|
- Your own Python code with stateful workers on a cluster: Ray Core
|
||||||
|
- ETL and SQL on structured data, or a JVM shop: PySpark
|
||||||
|
- MPI programs on HPC clusters and supercomputers: mpi4py
|
||||||
|
|
||||||
|
joblib is a [package for parallel computing and disk-based caching](https://joblib.readthedocs.io/en/stable/) that leaves your code as unmodified as possible. You [write a parallel for loop](https://joblib.readthedocs.io/en/stable/user_guide/parallel.html#common-usage) as a generator expression: `Parallel(n_jobs=2)(delayed(sqrt)(i ** 2) for i in range(10))`. By default, joblib runs the calls in separate worker processes. When one machine isn't enough, [switch its backend](https://joblib.readthedocs.io/en/stable/user_guide/parallel.html#setting-up-joblib-s-backend-with-parallel-config), and the same loop runs on a Dask, Ray, or Spark cluster. scikit-learn's [`n_jobs` runs on joblib](https://scikit-learn.org/stable/computing/parallelism.html) too, so it follows the backend you pick.
|
||||||
|
|
||||||
|
Dask [scales pandas, scikit-learn, and NumPy workflows](https://docs.dask.org/en/stable/why.html) with minimal rewriting. A Dask DataFrame is [a collection of pandas dataframes](https://docs.dask.org/en/stable/) on different computers, and Dask Arrays parallelize NumPy. Dask [runs without any setup](https://docs.dask.org/en/stable/deploying.html#local-machine) on your laptop. Its `LocalCluster` follows the same interface as every other Dask cluster manager, so you swap it out when you're ready to scale up.
|
||||||
|
|
||||||
|
Ray is a [unified framework for scaling AI and Python applications](https://docs.ray.io/en/latest/ray-overview/index.html). Its libraries each distribute one ML task: Ray Data, Train, Tune, Serve, and RLlib. Ray Core runs your own code: [decorate a function with `@ray.remote`](https://docs.ray.io/en/latest/ray-core/walkthrough.html), call it with `.remote()`, and fetch the result with `ray.get()`. For workers that keep state between calls, decorate a class the same way to get an actor. Ray runs on one machine with `ray.init()`, and on several nodes once you [deploy a Ray cluster](https://docs.ray.io/en/latest/cluster/getting-started.html). Ray Data [suits GPU workloads for deep learning inference](https://docs.ray.io/en/latest/data/comparisons.html#how-does-ray-data-compare-to-other-solutions-for-offline-inference) better than Spark does, but unlike Spark, it has no SQL interface.
|
||||||
|
|
||||||
|
PySpark is the Python API for Apache Spark, for large-scale data processing. Dask's own comparison suggests Spark [when you prefer SQL, run mostly JVM infrastructure, or want an all-in-one solution](https://docs.dask.org/en/stable/spark.html#reasons-you-might-choose-spark). PySpark's docs [recommend DataFrames over RDDs](https://spark.apache.org/docs/latest/api/python/index.html), so Spark builds the most efficient query for you. Write it in [SQL or the DataFrame API](https://spark.apache.org/docs/latest/api/python/user_guide/sql.html), whichever you think in, and switch between the two as you go. Installing PySpark with pip is [for local use or as a client](https://spark.apache.org/docs/latest/api/python/getting_started/install.html) that connects to a cluster, not for setting up the cluster itself. The `pyspark` package also needs Java.
|
||||||
|
|
||||||
|
mpi4py provides [Python bindings for MPI](https://mpi4py.readthedocs.io/en/stable/), the Message Passing Interface, so your Python code runs across workstations, clusters, and supercomputers. Lowercase methods like `comm.send` pass any picklable Python object, and uppercase ones like `comm.Send` [pass NumPy arrays the fast way](https://mpi4py.readthedocs.io/en/stable/tutorial.html). Run your script with `mpiexec -n 4 python -m mpi4py script.py`, so an unhandled exception [aborts the whole MPI run instead of deadlocking](https://mpi4py.readthedocs.io/en/stable/mpi4py.run.html#exceptions-and-deadlocks). mpi4py runs on an MPI implementation like MPICH or Open MPI, and in production its docs [recommend a custom-built or system-provided one](https://mpi4py.readthedocs.io/en/stable/install.html).
|
||||||
|
|
||||||
|
Start on one machine: parallelism [brings extra complexity and overhead](https://docs.dask.org/en/stable/best-practices.html#start-small), and often you don't need it. When you do scale out, collect results at the end, since Dask's `compute()` and Ray's `ray.get()` both block until the work finishes. Call [`compute()` once](https://docs.dask.org/en/stable/best-practices.html#avoid-calling-compute-repeatedly) rather than in a loop, and [`ray.get()` as late as possible](https://docs.ray.io/en/latest/ray-core/tips-for-first-time.html#tip-1-delay-ray-get).
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
Shipping to users without Python means converting a Python script to an executable, which PyInstaller does in one command. Pyarmor can obfuscate it first.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Standalone executable for most apps: PyInstaller
|
||||||
|
- Obfuscated scripts, bound to a machine or set to expire: Pyarmor
|
||||||
|
- Compiled code instead of bundled bytecode: Nuitka
|
||||||
|
- Tools for machines that already run Python: shiv
|
||||||
|
- Installers and Linux packages: cx_Freeze
|
||||||
|
|
||||||
|
PyInstaller [bundles your app and all its dependencies](https://pyinstaller.org/en/latest/) into one package, so users run it without installing Python or any modules. For most programs, that's [one short command](https://pyinstaller.org/en/latest/operating-mode.html): `pyinstaller myscript.py`.
|
||||||
|
|
||||||
|
Pyarmor [obfuscates Python scripts](https://pyarmor.readthedocs.io/en/latest/tutorial/getting-started.html), and can bind them to a machine or make them expire. `pyarmor gen foo.py` writes the obfuscated script to `dist/`. To ship an executable, Pyarmor [packs through PyInstaller](https://pyarmor.readthedocs.io/en/latest/tutorial/obfuscation.html#packing-obfuscated-scripts): `pyarmor gen --pack onefile foo.py`.
|
||||||
|
|
||||||
|
Without obfuscation, a PyInstaller bundle holds `.pyc` files that [could in principle be decompiled](https://pyinstaller.org/en/latest/operating-mode.html#hiding-the-source-code). Pyarmor is commercial: its free version is only for scripts that [won't make you a lot of money](https://pyarmor.readthedocs.io/en/latest/licenses.html#terms-of-use).
|
||||||
|
|
||||||
|
Nuitka is an optimizing Python compiler, and it [needs a C compiler](https://nuitka.net/user-documentation/user-manual.html). Its default mode needs Python on the machine, so [build in standalone mode to distribute](https://nuitka.net/user-documentation/tutorial-setup-and-build.html#distribute), and copy the resulting folder. Compiling protects your source code, but Nuitka's docs say constants stay readable unless you buy [Nuitka Commercial](https://nuitka.net/doc/commercial.html).
|
||||||
|
|
||||||
|
shiv builds [self-contained zipapps with all their dependencies included](https://shiv.readthedocs.io/en/latest/). `shiv -c hello -o hello .` packages your project the way `pip install .` would, with `-c` naming its console script. The result [depends on a pre-installed Python](https://packaging.python.org/en/latest/overview/#depending-on-a-pre-installed-python), which you can count on in your data centers and on developers' machines, but not on every user's computer.
|
||||||
|
|
||||||
|
cx_Freeze [freezes a script and its modules into a standalone executable](https://cx-freeze.readthedocs.io/en/latest/script.html), and also [builds installers and packages](https://cx-freeze.readthedocs.io/en/latest/): MSI for Windows, DMG for macOS, and deb, RPM, and AppImage for Linux. Put your options in `pyproject.toml` under `[tool.cxfreeze]` and build with [`cxfreeze build`](https://cx-freeze.readthedocs.io/en/latest/setup_script.html).
|
||||||
|
|
||||||
|
Get a folder build working before you switch to a single file, since problems are easier to diagnose in a folder. PyInstaller's docs [say so for one-folder mode](https://pyinstaller.org/en/latest/operating-mode.html#bundling-to-one-file), and Nuitka's [for standalone mode](https://nuitka.net/user-documentation/use-cases.html#standalone-program-distribution).
|
||||||
|
|
||||||
|
For an executable, plan to build on each OS you ship to. PyInstaller [isn't a cross-compiler](https://pyinstaller.org/en/latest/), cx_Freeze [only makes executables for the platform it runs on](https://cx-freeze.readthedocs.io/en/latest/faq.html#freezing-for-other-platforms), and Nuitka's docs suggest [Nuitka-Action](https://nuitka.net/user-documentation/use-cases.html#building-with-github-workflows) to build on all 3 OSes in GitHub workflows.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
Write only docstrings, and pdoc is all the Python documentation generator you need. Write guides too, and build the site with Sphinx or Material for MkDocs.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- API docs straight from docstrings, with no configuration: pdoc
|
||||||
|
- Handwritten docs plus an API reference, as HTML, PDF, and more: Sphinx
|
||||||
|
- A searchable Markdown site, with no HTML, CSS, or JavaScript to learn: Material for MkDocs
|
||||||
|
- Architecture diagrams as code, kept in version control: Diagrams
|
||||||
|
- A Material for MkDocs site, new or existing, on its team's own generator: Zensical
|
||||||
|
|
||||||
|
pdoc [aims to do one thing and do it well](https://pdoc.dev/docs/pdoc.html#what-is-pdoc): API documentation that follows your module hierarchy, with no configuration. Docstrings are Markdown, and it understands Google and numpydoc styles too. Run `pdoc ./demo.py` or `pdoc my_module_name` for a [preview in your browser that reloads](https://pdoc.dev/docs/pdoc.html#quickstart) when you edit the code, then `pdoc ./demo.py -o ./docs` to export the HTML. Its output is self-contained HTML, and for substantially more complex documentation needs, [pdoc's docs recommend Sphinx](https://pdoc.dev/docs/pdoc.html#limitations).
|
||||||
|
|
||||||
|
Sphinx [focuses on handwritten documentation](https://www.sphinx-doc.org/en/stable/usage/quickstart.html) and turns one set of source files into HTML, a PDF via LaTeX, man pages, and more. Its default markup is reStructuredText, and it can [read Markdown through MyST-Parser](https://www.sphinx-doc.org/en/stable/usage/markdown.html). Run `sphinx-quickstart` to set up a source directory with a `conf.py`, then list your pages in the root document's toctree. `make html` builds the site, and `make latexpdf` the PDF. To document your code, [autodoc](https://www.sphinx-doc.org/en/stable/usage/extensions/autodoc.html) pulls in its docstrings, which you mix with your handwritten pages.
|
||||||
|
|
||||||
|
Material for MkDocs is [a documentation framework on top of MkDocs](https://squidfunk.github.io/mkdocs-material/getting-started/), and `pip install mkdocs-material` installs MkDocs with it. You write Markdown and get a searchable static site, with [no HTML, CSS, or JavaScript to know](https://squidfunk.github.io/mkdocs-material/). Run `mkdocs new .`, then [set `site_name`, `site_url`, and `theme: name: material`](https://squidfunk.github.io/mkdocs-material/creating-your-site/#minimal-configuration) in `mkdocs.yml`, and `mkdocs serve` previews the site as you write. For reference docs from docstrings, try [mkdocstrings](https://squidfunk.github.io/mkdocs-material/alternatives/#sphinx) before switching to Sphinx: Material's alternatives page says it builds on MkDocs and adds Sphinx-like functionality.
|
||||||
|
|
||||||
|
Zensical is [a static site generator from the creators of Material for MkDocs](https://zensical.org/docs/get-started/), with the same batteries-included approach. After `pip install zensical`, [`zensical new .`](https://zensical.org/docs/create-your-site/) creates a `docs/` folder, a `zensical.toml` config, and a GitHub Actions workflow. `zensical serve` previews as you write, and `zensical build` writes the static site. It also [builds existing MkDocs projects without changes](https://zensical.org/docs/compatibility/mkdocs/) from their `mkdocs.yml`, and its classic theme variant keeps the Material for MkDocs look.
|
||||||
|
|
||||||
|
Diagrams [draws cloud system architecture in Python code](https://diagrams.mingrammer.com/), so you can track changes to a diagram in version control. It renders with Graphviz, so [install Graphviz first](https://diagrams.mingrammer.com/docs/getting-started/installation), then `pip install diagrams`. Describe the system in a `with Diagram("Web Service", show=False):` block and chain nodes with `>>`. Running `python diagram.py` saves it as a PNG in your working directory.
|
||||||
|
|
||||||
|
Install [Sphinx](https://www.sphinx-doc.org/en/stable/usage/installation.html), [Material for MkDocs](https://squidfunk.github.io/mkdocs-material/getting-started/#with-pip), or [Zensical](https://zensical.org/docs/get-started/#install-with-pip) into your project's virtual environment, as each one's docs recommend. The documentation generators here all write static HTML, so publish it from CI. All four document a GitHub Actions workflow that deploys to GitHub Pages: [Sphinx](https://www.sphinx-doc.org/en/stable/tutorial/deploying.html#publishing-your-html-documentation), [Material for MkDocs](https://squidfunk.github.io/mkdocs-material/publishing-your-site/), [Zensical](https://zensical.org/docs/publish-your-site/), and [pdoc](https://pdoc.dev/docs/pdoc.html#deploying-to-github-pages).
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
One function call sends mail through Gmail or another SMTP server with yagmail. This Python email library aims to make sending mail painless.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Gmail: yagmail, with an app password or OAuth2
|
||||||
|
- Another SMTP server: yagmail, pointed at its host
|
||||||
|
- An asyncio app: yagmail's async client
|
||||||
|
|
||||||
|
yagmail is a [wrapper around smtplib's SMTP connection](https://yagmail.readthedocs.io/en/latest/api.html#yagmail.Client) that connects to Gmail unless you pass another `host`. It builds the message for you: call `send()` with the recipients, a subject, and `contents`, a list whose strings it [reads as a local file, HTML, or text](https://yagmail.readthedocs.io/en/latest/usage.html#magical-contents). So one call sends text, HTML, and attachments. Wrap a string in `yagmail.raw` when it must stay plain text. In asyncio code, use its async client [as an async context manager](https://yagmail.readthedocs.io/en/latest/usage.html#starting-and-closing-connections).
|
||||||
|
|
||||||
|
Keep your password out of your script. Install `yagmail[all]` to get keyring, then [register your credentials once](https://yagmail.readthedocs.io/en/latest/setup.html#configuring-credentials) with `yagmail.register()`, and yagmail reads them from your system keyring. For Gmail, that password is an app password. For credentials you can revoke, [use OAuth2](https://yagmail.readthedocs.io/en/latest/setup.html#using-oauth2): whoever gets its token file can send mail, but nothing else.
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
Since uv both downloads Python and creates virtual environments, one Python environment manager does the jobs that pyenv and virtualenv split between them.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Python versions and virtual environments from one tool: uv
|
||||||
|
- A Python API for creating environments, or plugins: virtualenv
|
||||||
|
- Switching the `python` command between versions, and nothing more: pyenv
|
||||||
|
|
||||||
|
uv [manages Python versions](https://docs.astral.sh/uv/) as well as packages. When a command needs a Python you don't have, uv [downloads it for you](https://docs.astral.sh/uv/guides/install-python/), so you don't need Python installed to get started. A project's virtual environment lives in a `.venv` folder [inside the project, where editors can find it](https://docs.astral.sh/uv/concepts/projects/layout/#the-project-environment); keep it out of version control. Outside a project, `uv venv --python <version>` [creates an environment](https://docs.astral.sh/uv/pip/environments/) with that version, and downloads it if needed.
|
||||||
|
|
||||||
|
virtualenv [creates isolated Python environments](https://virtualenv.pypa.io/en/latest/), and a subset of it ships with Python as the venv module. Its docs place it [between venv and uv](https://virtualenv.pypa.io/en/latest/explanation.html#virtualenv-vs-venv-vs-uv): faster and more featureful than venv, and still pure Python. Pick it when you need plugins, or a Python API to create environments. It uses the Python it runs under unless you [pass `-p`](https://virtualenv.pypa.io/en/latest/how-to/usage.html#select-a-python-version) with another installed version. As a command-line tool, it [belongs in an isolated environment](https://virtualenv.pypa.io/en/latest/how-to/install.html), not your system Python.
|
||||||
|
|
||||||
|
pyenv does one job: it [lets you switch between multiple versions of Python](https://github.com/pyenv/pyenv), and leaves virtual environments [to you or its pyenv-virtualenv plugin](https://github.com/pyenv/pyenv#in-contrast-with-pythonbrew-and-pythonz-pyenv-does-not). Install it with the [automatic installer](https://github.com/pyenv/pyenv#1-automatic-installer-recommended) or, on macOS, Homebrew, then [set up your shell](https://github.com/pyenv/pyenv#b-set-up-your-shell-environment-for-pyenv) for it. [Choose the version](https://github.com/pyenv/pyenv#switch-between-python-versions) with `pyenv shell` for the session, `pyenv local` for a directory, or `pyenv global` for your user account. [Most versions are built from source](https://github.com/pyenv/pyenv#install-additional-python-versions) as you install them, so install Python's build dependencies first. pyenv [doesn't work on Windows](https://github.com/pyenv/pyenv#windows) outside the Windows Subsystem for Linux, so on Windows, use uv, which [supports Windows](https://docs.astral.sh/uv/).
|
||||||
|
|
||||||
|
Whichever tools you combine, name each project's Python in a `.python-version` file. `pyenv local` [writes that file](https://github.com/pyenv/pyenv#understanding-python-version-selection), and so does `uv python pin`, whose docs [recommend a plain version number](https://docs.astral.sh/uv/concepts/python-versions/#python-version-files) there so other tools can read it. virtualenv [uses the version pyenv selected](https://virtualenv.pypa.io/en/latest/how-to/usage.html#using-version-managers-pyenv-mise-asdf) with no extra configuration. Don't install packages into the Python that came with your operating system, which [often manages its packages itself](https://docs.astral.sh/uv/pip/environments/); a virtual environment keeps yours apart.
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
Out of the box, Odoo's CRM, eCommerce, and accounting apps run alone or combine into a full Python ERP. Anything they lack, you add as a module.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Trying Odoo, or customizing it without code: Odoo Online
|
||||||
|
- Writing your own modules: a source install
|
||||||
|
- Hosting your own modules in the cloud: Odoo.sh
|
||||||
|
- Free and open source: Odoo Community
|
||||||
|
- More features, with support and upgrades: Odoo Enterprise
|
||||||
|
|
||||||
|
Everything in Odoo starts and ends with modules, and [the main user-facing ones are flagged as Apps](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/01_architecture.html#odoo-modules). Odoo Enterprise is [extra modules installed on top of Community](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/01_architecture.html#odoo-editions). [Community is free and open source under the LGPL](https://www.odoo.com/documentation/latest/administration.html#editions). Enterprise is shared source, and its license [ties running it to an Enterprise subscription](https://www.odoo.com/documentation/latest/legal/licenses.html).
|
||||||
|
|
||||||
|
[Odoo Online](https://www.odoo.com/documentation/latest/administration/odoo_online.html) runs in your browser with nothing to install, and handles customizations that need no code. Your own modules need Odoo.sh or your own server. Odoo.sh is Odoo's official cloud platform, and it [builds your modules from a GitHub repository](https://www.odoo.com/documentation/latest/administration/odoo_sh/create_module.html).
|
||||||
|
|
||||||
|
To develop modules, Odoo's developer docs prefer [a source install](https://www.odoo.com/documentation/latest/developer/tutorials/setup_guide.html), which runs Odoo straight from its code. Odoo's logic is [written in Python, and it stores data only in PostgreSQL](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/01_architecture.html#multitier-application). Keep your modules in a directory of their own, and [start the server with `odoo-bin`](https://www.odoo.com/documentation/latest/administration/on_premise/source.html#running-odoo), adding that directory to `--addons-path`. Then work through [Server framework 101](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101.html), which builds one module chapter by chapter.
|
||||||
|
|
||||||
|
A module can [add new business logic or change what's already there](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/01_architecture.html#odoo-modules). To change a standard model, extend it from your own module: [model inheritance](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/12_inheritance.html#model-inheritance) adds fields and overrides methods on a model another module defines. Screens work the same way: [view inheritance](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/12_inheritance.html#view-inheritance) applies your extension views on top of the originals instead of overwriting them.
|
||||||
|
|
||||||
|
Before you write a module, [check whether Odoo already covers the case](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/02_newapp.html). A database with custom modules [can't be upgraded until they're ready for the new version](https://www.odoo.com/documentation/latest/administration/upgrade.html), so [cut what duplicates the standard modules](https://www.odoo.com/documentation/latest/developer/howtos/upgrade_custom_db.html#step-1-stop-the-developments).
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
PDF, Word, and Excel files open with pypdf, python-docx, and openpyxl, a Python file format library for each. MarkItDown turns all three into Markdown.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Splitting, merging, and reading PDFs: pypdf
|
||||||
|
- Word documents: python-docx
|
||||||
|
- Excel: openpyxl to read or edit, XlsxWriter for new reports
|
||||||
|
- Documents to Markdown for an LLM: MarkItDown
|
||||||
|
- ELF binaries and DWARF debug info: pyelftools
|
||||||
|
- One table exported to CSV, JSON, Excel, and more: Tablib
|
||||||
|
- Scanned PDFs, tables, and complex layouts: Docling
|
||||||
|
- PowerPoint decks: python-pptx
|
||||||
|
- New PDFs drawn from Python code: ReportLab
|
||||||
|
- PDF text with its position and font: pdfminer.six
|
||||||
|
- PDFs from HTML and CSS: WeasyPrint
|
||||||
|
- Markdown to HTML: markdown-it-py for CommonMark, Python-Markdown for its extensions, Mistune for speed
|
||||||
|
- Config files: tomllib for TOML, PyYAML for YAML
|
||||||
|
|
||||||
|
pypdf works on PDFs that already exist: it [splits, merges, crops, and transforms pages](https://pypdf.readthedocs.io/en/stable/), adds passwords, and pulls out text and metadata. It's [pure Python with no C dependency](https://pypdf.readthedocs.io/en/stable/meta/comparisons.html), and it doesn't create PDFs. pypdf [isn't OCR software](https://pypdf.readthedocs.io/en/stable/user/extract-text.html), so run scanned pages through OCR instead.
|
||||||
|
|
||||||
|
python-docx [only edits existing documents](https://python-docx.readthedocs.io/en/latest/user/documents.html): `Document()` opens a built-in template with no content. Start from your own .docx instead, so its styles, headers, and footers carry over.
|
||||||
|
|
||||||
|
openpyxl reads and writes Excel files. For big workbooks, open them with `read_only=True` or create them with `write_only=True`, which [keep memory near constant](https://openpyxl.readthedocs.io/en/stable/optimized.html). A cell with a formula loads the formula; pass [`data_only=True`](https://openpyxl.readthedocs.io/en/stable/tutorial.html) to get the value Excel last calculated.
|
||||||
|
|
||||||
|
XlsxWriter [only writes new files](https://xlsxwriter.readthedocs.io/introduction.html) and can't read or modify existing ones, but it supports more Excel features than the alternatives. Open the workbook [in a `with` block](https://xlsxwriter.readthedocs.io/workbook.html) so it gets closed and saved. For large files, turn on [`constant_memory`](https://xlsxwriter.readthedocs.io/working_with_memory.html) and write the rows in order. From pandas, pass [`engine='xlsxwriter'`](https://xlsxwriter.readthedocs.io/working_with_pandas.html) to `pd.ExcelWriter`.
|
||||||
|
|
||||||
|
MarkItDown converts files to Markdown [for LLMs and text analysis](https://github.com/microsoft/markitdown), not for high-fidelity conversions that people read. Install `markitdown[all]`, or only the extras for your formats, like `markitdown[pdf, docx, pptx]`.
|
||||||
|
|
||||||
|
Docling [understands PDF layout](https://docling-project.github.io/docling/): reading order, tables, and formulas. It also runs OCR on scanned pages. Its models [run locally and send no data out](https://docling-project.github.io/docling/usage/advanced_options/#using-remote-services) unless you turn remote services on. Each file becomes a DoclingDocument, which you export to Markdown or [split into chunks](https://docling-project.github.io/docling/concepts/chunking/) for an embedding model.
|
||||||
|
|
||||||
|
python-pptx builds decks from data, like a database query or analytics output, and [doesn't need PowerPoint installed](https://python-pptx.readthedocs.io/en/latest/). It [only edits existing presentations](https://python-pptx.readthedocs.io/en/latest/user/presentations.html), so start from your own deck: its theme, slide master, and slide layouts set how the slides look. Add each slide from one of those layouts, picked by [its index in your deck](https://python-pptx.readthedocs.io/en/latest/user/slides.html).
|
||||||
|
|
||||||
|
ReportLab draws new PDFs from Python code. Learn it on [`pdfgen`, its lowest-level interface](https://docs.reportlab.com/developerfaqs/), then build multi-page documents with [Platypus](https://docs.reportlab.com/reportlab/userguide/ch5_platypus/). Platypus lets you keep paragraph styles and page layouts in one shared file, so restyling takes a few lines.
|
||||||
|
|
||||||
|
pdfminer.six [focuses on text](https://github.com/pdfminer/pdfminer.six): it gets each piece of text with its exact location, font, and color. Start with `extract_text()` from its [high-level API](https://pdfminersix.readthedocs.io/en/latest/tutorial/highlevel.html). A PDF [stores only characters and their positions](https://pdfminersix.readthedocs.io/en/latest/topic/converting_pdf_to_text.html), so pdfminer.six guesses words, lines, and paragraphs from the layout. Tune those guesses with `LAParams`.
|
||||||
|
|
||||||
|
WeasyPrint turns HTML and CSS into PDFs, like [reports, invoices, and tickets](https://doc.courtbouillon.org/weasyprint/stable/). It runs [no JavaScript](https://doc.courtbouillon.org/weasyprint/stable/going_further.html). Set page size and margins [with the CSS `@page` rule](https://doc.courtbouillon.org/weasyprint/stable/common_use_cases.html).
|
||||||
|
|
||||||
|
markdown-it-py [follows the CommonMark spec](https://markdown-it-py.readthedocs.io/en/latest/) and takes plugins for more syntax. For content your users submit, use the [`js-default` preset](https://markdown-it-py.readthedocs.io/en/latest/security.html), since the default settings aren't safe for it.
|
||||||
|
|
||||||
|
Python-Markdown [isn't a CommonMark implementation](https://python-markdown.github.io/): it follows the original Markdown syntax and has an extension API. It [doesn't sanitize its HTML output](https://python-markdown.github.io/sanitization/), so sanitize it yourself when the input is untrusted.
|
||||||
|
|
||||||
|
Mistune is [fast and has no dependencies](https://mistune.lepture.com/en/latest/). For untrusted input, build the parser with [`mistune.create_markdown()`](https://mistune.lepture.com/en/latest/guide.html), which escapes HTML tags, since `mistune.html()` doesn't.
|
||||||
|
|
||||||
|
tomllib [only reads TOML](https://docs.python.org/3/library/tomllib.html), from a file opened in binary mode.
|
||||||
|
|
||||||
|
Tablib holds one dataset and exports it to many formats; Excel, YAML, and pandas [are optional extras](https://tablib.readthedocs.io/en/stable/formats.html), like `tablib[xlsx]`.
|
||||||
|
|
||||||
|
pyelftools is [pure Python with no dependencies](https://github.com/eliben/pyelftools); start from its [`ELFFile` class](https://github.com/eliben/pyelftools/blob/main/doc/user-guide.md) and stay on the high-level API.
|
||||||
|
|
||||||
|
Treat every file you didn't create as untrusted. With PyYAML, call [`yaml.safe_load()`](https://pyyaml.org/wiki/PyYAMLDocumentation#loading-yaml), never `yaml.load()`, which can run any Python function. Install [defusedxml](https://openpyxl.readthedocs.io/en/stable/#security) next to openpyxl to guard against XML attacks like billion laughs. [Catch pypdf's exceptions](https://pypdf.readthedocs.io/en/stable/user/security.html) yourself, so a broken PDF can't crash your service. For MarkItDown, call [`convert_local()` or `convert_stream()`](https://github.com/microsoft/markitdown#security-considerations) instead of `convert()`, which also fetches remote URIs. Cap Docling's input with [`max_num_pages` and `max_file_size`](https://docling-project.github.io/docling/usage/advanced_options/#impose-limits-on-the-document-size). Run WeasyPrint on untrusted HTML [as a user with limited access](https://doc.courtbouillon.org/weasyprint/stable/first_steps.html#security), with a URL fetcher that blocks local files. With tomllib, [limit the size](https://docs.python.org/3/library/tomllib.html) of the data you parse.
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
Without installing anything, two Python file manipulation libraries cover paths and file types by name: pathlib and mimetypes. python-magic reads the bytes.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Building and joining paths, and listing directories: pathlib
|
||||||
|
- A MIME type from a file name or URL: mimetypes
|
||||||
|
- A file's type from its content, like the Unix `file` command: python-magic
|
||||||
|
- Rerunning or reloading code when files change: watchfiles
|
||||||
|
- Your own handler for each created, modified, or moved file: watchdog
|
||||||
|
|
||||||
|
If you aren't sure which pathlib class you need, the docs say [`Path` is most likely it](https://docs.python.org/3/library/pathlib.html): it makes a concrete path for the platform your code runs on. [Join paths with the `/` operator](https://docs.python.org/3/library/pathlib.html#basic-use), and list a directory with `iterdir()` or `glob()`.
|
||||||
|
|
||||||
|
mimetypes [maps a file name's extension to a MIME type](https://docs.python.org/3/library/mimetypes.html), and a MIME type back to extensions. A type guess returns 2 values: a type for the Content-Type header, and an encoding like gzip for the Content-Encoding header. The type is [`None` for a missing or unknown suffix](https://docs.python.org/3/library/mimetypes.html#mimetypes.guess_type).
|
||||||
|
|
||||||
|
python-magic wraps libmagic, which [identifies file types by checking their headers](https://github.com/ahupp/python-magic), the way the Unix `file` command does. It's a thin wrapper, so install libmagic too, with `apt-get install libmagic1` or `brew install libmagic`. Call `magic.from_file(path, mime=True)` for a MIME type, or `magic.from_buffer()` on bytes you already have.
|
||||||
|
|
||||||
|
watchfiles is built for [file watching and code reload](https://watchfiles.helpmanual.io/), and its Rust core groups changes into batches instead of firing once per file. `watch()` is a generator that [yields sets of changes](https://watchfiles.helpmanual.io/api/watch/#watchfiles.watch), and `awatch()` is the async version. To restart code when files change, use [`run_process()`](https://watchfiles.helpmanual.io/api/run_process/#watchfiles.run_process) with a function or a command. From a shell, the [`watchfiles` CLI](https://watchfiles.helpmanual.io/cli/) does the same: `watchfiles --filter python 'pytest --lf' src tests`.
|
||||||
|
|
||||||
|
watchdog gives you [an API and a shell tool](https://python-watchdog.readthedocs.io/en/latest/) for monitoring directories. Its [quickstart](https://python-watchdog.readthedocs.io/en/latest/quickstart.html) has you subclass `FileSystemEventHandler`, override methods like `on_created()` and `on_modified()`, schedule the handler on an `Observer`, and start that thread. An observer skips subdirectories unless you pass `recursive=True` to `schedule()`.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
Beyond functools, install more-itertools, as the itertools docs suggest. toolz is a fuller Python functional programming library, and returns adds typed errors.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Partial application, decorators, and caching: functools
|
||||||
|
- More iterator tools, the itertools recipes included: more-itertools
|
||||||
|
- Composing functions into pipelines, currying, and dict helpers: toolz, or cytoolz for speed
|
||||||
|
- Everyday helpers for collections, decorators, retries, and debugging: funcy
|
||||||
|
- Errors and missing values as typed containers checked by mypy: returns
|
||||||
|
|
||||||
|
functools is the standard library's module [for higher-order functions](https://docs.python.org/3/library/functools.html), functions that act on or return other functions. Python's Functional Programming HOWTO calls `partial()` [the most useful tool in the module](https://docs.python.org/3/howto/functional.html#the-functools-module): it fills in some of a function's arguments and gives you a new function. When you write a decorator, wrap its inner function with [`wraps`](https://docs.python.org/3/library/functools.html#functools.wraps), so the decorated function keeps its name and docstring. The same HOWTO finds many uses of `reduce()` [clearer as a `for` loop](https://docs.python.org/3/howto/functional.html#small-functions-and-the-lambda-expression).
|
||||||
|
|
||||||
|
The itertools docs point you to more-itertools for [their recipes and many more](https://docs.python.org/3/library/itertools.html#itertools-recipes). It collects [building blocks beyond itertools](https://more-itertools.readthedocs.io/en/stable/), for grouping, windowing, lookahead, and more. The itertools recipes sit in its top-level package, so `from more_itertools import flatten` works.
|
||||||
|
|
||||||
|
toolz [extends itertools and functools](https://toolz.readthedocs.io/en/latest/) with functions that are composable, pure, and lazy, and its API [follows Clojure's standard library](https://toolz.readthedocs.io/en/latest/heritage.html). Each function takes and returns only iterables, dictionaries, and functions, so they [compose to solve your own problems](https://toolz.readthedocs.io/en/latest/composition.html). [`pipe`](https://toolz.readthedocs.io/en/latest/api.html#toolz.functoolz.pipe) runs a value through a sequence of functions, like pipes in Unix. Stick with `partial` at first, and once it shows up several times in your code, [switch to the `toolz.curried` namespace](https://toolz.readthedocs.io/en/latest/curry.html#curry). toolz is a general-purpose library, and for data analytics its docs say [a library built for it](https://toolz.readthedocs.io/en/latest/streaming-analytics.html#disclaimer) may serve you better. [cytoolz](https://github.com/pytoolz/cytoolz) implements the same API in Cython, as a drop-in replacement when you need more speed.
|
||||||
|
|
||||||
|
funcy is a collection of functional tools [focused on practicality](https://github.com/Suor/funcy), inspired by Clojure and underscore. Next to sequence tools, it has [collection functions that keep the type](https://funcy.readthedocs.io/en/stable/overview.html) of a dict or set. It also has control flow helpers, like `@retry` and `silent`, and debugging helpers, like `tap` and `log_calls`. Many of its functions take a regex, a mapping, or a set [where you'd pass a function](https://funcy.readthedocs.io/en/stable/extended_fns.html#extended-function-semantics).
|
||||||
|
|
||||||
|
returns puts results in typed containers: [`Maybe` for None and `Result` for exceptions](https://returns.readthedocs.io/en/latest/pages/quickstart.html#why), plus `IO` for impure code and `Future` for async code. Its docs [really recommend mypy](https://returns.readthedocs.io/en/latest/pages/quickstart.html#typechecking-and-other-integrations), and typing [only works correctly with its mypy plugin](https://returns.readthedocs.io/en/latest/pages/result.html). So it fits projects that check types with mypy. Turn functions that raise into ones that return a `Result` with [`@safe`](https://returns.readthedocs.io/en/latest/pages/result.html#safe), and chain the steps with [`flow`](https://returns.readthedocs.io/en/latest/pages/pipeline.html#flow), which its docs call the recommended way to write code with returns.
|
||||||
|
|
||||||
|
Mix functional style with the rest of your code: Python's HOWTO says functional-style programs usually [give a functional-appearing interface](https://docs.python.org/3/howto/functional.html) and use non-functional features inside. For a plain map or filter, toolz's own docs call comprehensions [more Pythonic](https://toolz.readthedocs.io/en/latest/streaming-analytics.html). Learn a core set of functions; toolz says [about a dozen covers most tasks](https://toolz.readthedocs.io/en/latest/control.html), and the right word only helps when your readers know it too.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
A 2D game needs a loop: write your own with pygame-ce, the Python game development library, or let Arcade run it. Panda3D does 3D, and Ren'Py visual novels.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- 2D game where you write the game loop: pygame-ce
|
||||||
|
- 2D game with a ready-made loop and physics engines: Arcade
|
||||||
|
- 3D game: Panda3D
|
||||||
|
- Visual novel or life simulation game: Ren'Py
|
||||||
|
- Windowing, OpenGL graphics, and sound with no other dependencies: pyglet
|
||||||
|
|
||||||
|
pygame-ce is the community edition of pygame, a fork by its former core developers that aims for more frequent releases. Your code still says `import pygame`, so if pygame is already installed, [uninstall it first](https://github.com/pygame-community/pygame-ce/wiki/Installing-pygame%E2%80%90ce), then run `pip install pygame-ce` in a virtual environment. Its [quick start](https://pyga.me/docs/) gives you full control of the game loop: handle events, draw the frame, `flip()` the display, and cap the frame rate with `clock.tick()`. Convert each image once after you load it, with [`convert()`, or `convert_alpha()` if it has transparency](https://pyga.me/docs/ref/surface.html#pygame.Surface.convert), so it blits fast. It's under the LGPL, and its README says [closed-source and commercial games are fine](https://github.com/pygame-community/pygame-ce).
|
||||||
|
|
||||||
|
Arcade is [an easy-to-learn library for 2D games](https://github.com/pythonarcade/arcade), built on pyglet and OpenGL, and meant for beginning programmers too. Start with the [Platformer Tutorial](https://api.arcade.academy/en/stable/tutorials/platform_tutorial/step_01.html): you subclass `arcade.Window`, draw in `on_draw()`, and `arcade.run()` runs the loop until the window closes. For movement and collisions, it comes with [physics engines](https://api.arcade.academy/en/stable/api_docs/api/physics_engines.html) for top-down and platformer games. Its code is MIT, and its [built-in assets need no attribution](https://api.arcade.academy/en/stable/), so you can ship them in a commercial game.
|
||||||
|
|
||||||
|
Panda3D is [a 3D engine written in C++ with Python bindings](https://docs.panda3d.org/latest/python/introduction/index), and its manual says it's a tool for skilled programmers, not beginners. Install it with [`pip install panda3d`](https://github.com/panda3d/panda3d), subclass [`ShowBase`](https://docs.panda3d.org/latest/python/introduction/tutorial/starting-panda3d), and call `run()`, which holds the main loop. Its [`build_apps` tool](https://docs.panda3d.org/latest/python/distribution/index) builds self-contained executables for Windows, Linux, and macOS without needing each system. It's BSD, free for commercial games.
|
||||||
|
|
||||||
|
Ren'Py is a [visual novel engine](https://www.renpy.org/) with its own script language, for stories that run on computers and mobile devices. It isn't a pip package: [download Ren'Py and run its launcher](https://www.renpy.org/doc/html/quickstart.html), create a project there, and write your story in `script.rpy`. Python works inside the scripts, and [third-party pure-Python packages](https://www.renpy.org/doc/html/python.html#first-and-third-party-python-modules-and-packages) go in `game/python-packages`. Ship with [Build Distributions](https://www.renpy.org/doc/html/build.html) in the launcher, which also builds a package for itch.io and Steam. Most of Ren'Py is MIT, but some parts are LGPL, so [distribute your game in a way that satisfies the LGPL](https://www.renpy.org/doc/html/license.html).
|
||||||
|
|
||||||
|
pyglet is a [windowing and multimedia library with no external dependencies](https://pyglet.readthedocs.io/en/latest/), written in pure Python: windows, input, OpenGL graphics, images, video, and sound. Start with [Writing a pyglet application](https://pyglet.readthedocs.io/en/latest/programming_guide/quickstart.html), which attaches handlers with `@window.event` and calls `pyglet.app.run()`. Draw through a `Batch`, since the docs say [you always want batched rendering](https://pyglet.readthedocs.io/en/latest/programming_guide/shapes.html) for performance. It's under the BSD license.
|
||||||
|
|
||||||
|
Move things by the time since the last frame, so your game runs at the same speed at any frame rate. pygame-ce's quick start gets it in seconds by [dividing `clock.tick()` by 1000](https://pyga.me/docs/), pyglet passes it as `dt` to [scheduled functions](https://pyglet.readthedocs.io/en/latest/programming_guide/time.html#sprite-movement-techniques), and Arcade passes it as `delta_time` to [`on_update()`](https://api.arcade.academy/en/stable/api_docs/api/window.html#arcade.Window.on_update).
|
||||||
@@ -0,0 +1,20 @@
|
|||||||
|
Addresses become coordinates with geopy, whose Python geolocation API covers many geocoding services. GeoPandas analyzes map data, and GeoDjango serves it.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Addresses to coordinates, and coordinates back to addresses: geopy
|
||||||
|
- The distance between two points: geopy
|
||||||
|
- Geographic data in tables and files, like shapefiles and GeoJSON: GeoPandas
|
||||||
|
- Web apps that store and query geographic data: GeoDjango
|
||||||
|
- A visitor's country or city from their IP address, in a Django project: GeoDjango
|
||||||
|
- Building, encoding, and validating GeoJSON objects by hand: geojson
|
||||||
|
|
||||||
|
geopy is [a client for geocoding services, not a service itself](https://geopy.readthedocs.io/en/stable/#geopy-is-not-a-service), and each service has its own terms of use, quotas, and pricing. Every geocoder has a `geocode()` method that turns an address into a location, and most also have `reverse()` for the other way around. For OpenStreetMap's Nominatim, [set a `user_agent` that names your app](https://geopy.readthedocs.io/en/stable/#geopy.geocoders.Nominatim) and follow [its usage policy](https://operations.osmfoundation.org/policies/nominatim/). The policy bans heavy use, auto-complete, and systematic queries. It also asks you to cache results and show attribution, and it discourages bulk geocoding. To geocode a DataFrame, [wrap the call in `RateLimiter`](https://geopy.readthedocs.io/en/stable/#usage-with-pandas), which adds delays between requests and retries failed ones. Check first that your service allows bulk requests at all. For distances, `geopy.distance.distance` [computes the geodesic distance](https://geopy.readthedocs.io/en/stable/#module-geopy.distance) between two points.
|
||||||
|
|
||||||
|
GeoPandas adds geometry columns to pandas, so you can [do in Python what would otherwise take a spatial database such as PostGIS](https://geopandas.org/en/stable/). Load a shapefile, GeoJSON, or GeoPackage with `read_file()`, which [detects the file type and returns a GeoDataFrame](https://geopandas.org/en/stable/getting_started/introduction.html), and write it back with `to_file()`. For a table of latitudes and longitudes, [build the points with `points_from_xy()`](https://geopandas.org/en/stable/gallery/create_geopandas_from_pandas.html) and set `crs="EPSG:4326"`. To geocode a column of addresses, GeoPandas [calls geopy for you](https://geopandas.org/en/stable/docs/user_guide/geocoding.html), so the geocoding service's terms still apply.
|
||||||
|
|
||||||
|
GeoDjango is [a contrib module that ships with Django](https://docs.djangoproject.com/en/stable/ref/contrib/gis/tutorial/). It adds model fields for geometries, spatial queries to the ORM, and geometry editing to the admin. Its tutorial assumes you already know Django. Run it on PostGIS, which its docs recommend as [the most mature and feature-rich open source spatial database](https://docs.djangoproject.com/en/stable/ref/contrib/gis/install/#spatial-database). To load a shapefile, [`ogrinspect` writes the model and a `LayerMapping` imports the rows](https://docs.djangoproject.com/en/stable/ref/contrib/gis/tutorial/#importing-spatial-data). For IP geolocation, [GeoIP2](https://docs.djangoproject.com/en/stable/ref/contrib/gis/geoip2/) looks up a country or city in a MaxMind or DB-IP database file you download.
|
||||||
|
|
||||||
|
geojson has [a class for every object in the GeoJSON spec](https://github.com/jazzband/geojson), and `geojson.dumps()` and `geojson.loads()` [wrap the standard json functions](https://github.com/jazzband/geojson#geojson-encodingdecoding) to encode and decode them. Check an object with [its `is_valid` property and `errors()` method](https://github.com/jazzband/geojson#validation). To encode your own classes the same way, [give them a `__geo_interface__`](https://github.com/jazzband/geojson#custom-classes), which GeoPandas' `from_features()` [also accepts](https://geopandas.org/en/stable/docs/reference/api/geopandas.GeoDataFrame.from_features.html).
|
||||||
|
|
||||||
|
Latitude and longitude are angles, so measure distance and area in a projected coordinate system, in meters or feet. In GeoPandas, you [always need one](https://geopandas.org/en/stable/getting_started/introduction.html#Projections) for those operations, so reproject with `to_crs()` first. GeoDjango calls choosing a geometry field's SRID [an important decision](https://docs.djangoproject.com/en/stable/ref/contrib/gis/model-api/#selecting-an-srid), since projected systems ease distance calculations. geopy's `distance` works on an ellipsoidal model of the earth, so it takes latitudes and longitudes as they are.
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
Among Python GUI libraries, PySide6 is the default for desktop apps and tkinter for small tools. For a UI in the browser, pick NiceGUI.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Full desktop app: PySide6, or PyQt6 if your app can be GPL or you buy a license
|
||||||
|
- Small tool without third-party packages: tkinter
|
||||||
|
- Modern look for a tkinter app: CustomTkinter
|
||||||
|
- tkinter layout drawn in Figma: Tkinter Designer
|
||||||
|
- Native widgets on Windows, macOS, and Linux: wxPython
|
||||||
|
- GNOME app on Linux: PyGObject
|
||||||
|
- GPU-rendered tools for your scripts: Dear PyGui
|
||||||
|
- Multi-touch apps on Android and iOS: Kivy
|
||||||
|
- Native widgets on desktop and mobile: Toga
|
||||||
|
- One codebase for web, desktop, and mobile: Flet
|
||||||
|
- Dashboards and web UIs: NiceGUI
|
||||||
|
- HTML/JavaScript frontend in a desktop window: pywebview
|
||||||
|
- GUI for an existing argparse script: Gooey
|
||||||
|
|
||||||
|
PySide6 is Qt for Python, the [official Python bindings for Qt](https://doc.qt.io/qtforpython-6/), under the LGPL, the GPL, or a commercial license. Qt's docs [recommend a virtual environment](https://doc.qt.io/qtforpython-6/gettingstarted.html) over installing it into your system Python. Ship it with [pyside6-deploy](https://doc.qt.io/qtforpython-6/deployment/index.html).
|
||||||
|
|
||||||
|
PyQt6 wraps the same Qt, but Riverbank [licenses it under the GPL or a commercial license, not the LGPL](https://www.riverbankcomputing.com/software/pyqt/). So a closed-source app needs a commercial PyQt6 license, while PySide6 can stay on the LGPL.
|
||||||
|
|
||||||
|
tkinter is the standard Python interface to Tcl/Tk. Python's docs recommend the [themed tkinter.ttk widgets](https://docs.python.org/3/library/tkinter.ttk.html), which follow the platform's native theme, over the classic ones most online docs still use.
|
||||||
|
|
||||||
|
NiceGUI runs a web server and shows your UI in the browser, which suits dashboards, micro web apps, and robotics projects. Pass `native=True` to `ui.run()` to [open it in a desktop window](https://nicegui.io/documentation/section_configuration_deployment) instead, or bundle it into an executable with nicegui-pack.
|
||||||
|
|
||||||
|
Kivy runs the same code on Android, iOS, Linux, macOS, and Windows. Declare the widget tree in the [KV language](https://kivy.org/doc/stable/guide/lang.html) to keep the UI apart from your logic, and build Android packages with [Buildozer](https://kivy.org/doc/stable/guide/packaging-android.html).
|
||||||
|
|
||||||
|
Toga [uses native system widgets, not themes](https://toga.beeware.org/en/stable/about/philosophy/), so a Toga app is a native app on each platform. Start with the [BeeWare tutorial](https://tutorial.beeware.org/), which packages your app with Briefcase.
|
||||||
|
|
||||||
|
Flet builds web, desktop, and mobile apps from one Python codebase, [without HTML, CSS, or JavaScript](https://flet.dev/docs/). Package it for each platform with [flet build](https://flet.dev/docs/publish/).
|
||||||
|
|
||||||
|
Whatever you pick, keep slow work out of event handlers, or the window freezes. The [tkinter docs](https://docs.python.org/3/library/tkinter.html) say to break it into smaller pieces with timers or run it in another thread, and Qt's docs suggest threads for the same reason.
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
Not every Python hardware library needs a board: pynput drives your keyboard and mouse, Bleak your Bluetooth LE devices. Jumpstarter automates hardware tests.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Controlling or monitoring the keyboard and mouse: pynput
|
||||||
|
- Bluetooth Low Energy devices, like sensors: Bleak
|
||||||
|
- Automated tests on real or virtual hardware: Jumpstarter
|
||||||
|
|
||||||
|
pynput [controls and monitors input devices](https://pynput.readthedocs.io/en/latest/): the mouse and the keyboard. To send input, create a `Controller` and [call its methods](https://pynput.readthedocs.io/en/latest/keyboard.html#controlling-the-keyboard), like `press()`, `release()`, or `type()` for a whole string. To react to input, open a `Listener` with your callbacks in a `with` block and [call `join()`](https://pynput.readthedocs.io/en/latest/keyboard.html#monitoring-the-keyboard). In a GUI app with its own main loop, call `start()` instead, so your code keeps running.
|
||||||
|
|
||||||
|
Bleak is a [GATT client](https://bleak.readthedocs.io/en/latest/): it connects to Bluetooth Low Energy devices, like sensors, through one asynchronous, cross-platform API. Connect in an `async with BleakClient(...)` block and start your program with `asyncio.run()`. That's [the recommended way](https://bleak.readthedocs.io/en/latest/api/client.html#connecting-and-disconnecting), and the device disconnects even when your program is interrupted or raises. Scan the same way, in an [`async with BleakScanner(...)` block](https://bleak.readthedocs.io/en/latest/api/scanner.html#starting-and-stopping).
|
||||||
|
|
||||||
|
Jumpstarter is an [open source framework for hardware-in-the-loop testing](https://jumpstarter.dev/main/introduction/index.html#introduction) on physical hardware and virtual devices. A person in `jmp shell`, a pytest script, and a CI pipeline all use the same APIs. [Local mode](https://jumpstarter.dev/main/introduction/index.html#local-mode) needs no Kubernetes and suits one developer with the hardware at hand. [Distributed mode](https://jumpstarter.dev/main/introduction/index.html#distributed-mode) runs a Kubernetes-based controller that leases devices, so teams can share them, including from CI. Write tests on the `JumpstarterTest` base class from jumpstarter-testing, which [handles the connection](https://jumpstarter.dev/main/getting-started/guides/examples/testing.html#the-jumpstartertest-base-class) for you. Set a `selector` for the device you need: the class connects from inside `jmp shell`, or leases a matching device outside it.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
Scraped HTML, however broken, parses in Beautiful Soup on lxml's parser. JustHTML packs a sanitizer into a Python HTML library and runs it by default.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Pulling data out of scraped pages: Beautiful Soup, on lxml's parser
|
||||||
|
- Speed, XPath, or XML documents: lxml
|
||||||
|
- Sanitizing HTML your users submit: JustHTML
|
||||||
|
- XML you'd rather handle like JSON: xmltodict
|
||||||
|
- Escaping text you put into HTML: MarkupSafe
|
||||||
|
|
||||||
|
Beautiful Soup is for [pulling data out of HTML and XML files](https://www.crummy.com/software/BeautifulSoup/bs4/doc/). It sits on a parser you pick and gives you one way to navigate, search, and change the tree. Its docs [recommend lxml as that parser](https://www.crummy.com/software/BeautifulSoup/bs4/doc/#installing-a-parser) for speed. Name the parser in the constructor, as in `BeautifulSoup(markup, "lxml")`. [Different parsers build different trees](https://www.crummy.com/software/BeautifulSoup/bs4/doc/#differences-between-parsers) from the same broken page, and your script should parse it the same way on every machine. The docs also say when to skip Beautiful Soup: [work directly atop lxml](https://www.crummy.com/software/BeautifulSoup/bs4/doc/#improving-performance) when computer time costs more than programmer time, and [parse with lxml](https://www.crummy.com/software/BeautifulSoup/bs4/doc/#css-selectors-through-the-css-property) when CSS selectors are all you need.
|
||||||
|
|
||||||
|
lxml is [a Pythonic binding for libxml2 and libxslt](https://lxml.de/): the speed and XML features of those C libraries, with an API mostly compatible with ElementTree's. For HTML, use [lxml.html](https://lxml.de/lxmlhtml.html), which adds HTML-specific methods to lxml's elements, and every element takes `.xpath()`. On broken pages, lxml's docs say you often need only [Beautiful Soup's encoding detection](https://lxml.de/lxmlhtml.html#really-broken-pages). Leave the rest to lxml's own parser, which is several times faster. XPath has the same injection problem as SQL. When a value comes from outside, [pass it as an XPath variable](https://lxml.de/FAQ.html#how-do-i-use-lxml-safely-as-a-web-service-endpoint) instead of formatting it into the expression.
|
||||||
|
|
||||||
|
JustHTML [parses HTML like a browser](https://github.com/EmilStenstrom/justhtml), including broken markup. `JustHTML(html)` [sanitizes by default](https://emilstenstrom.github.io/justhtml/sanitization.html) against a strict allowlist. The output is safe in a page body, but [not automatically safe inside a `<script>` tag or an attribute](https://emilstenstrom.github.io/justhtml/sanitization.html#important-context-is-king); JustHTML has separate escaping helpers for those. It's pure Python, with no C extension to install. For terabytes of trusted HTML, its README says to use a C or Rust parser like lxml.
|
||||||
|
|
||||||
|
xmltodict [makes working with XML feel like working with JSON](https://github.com/martinblech/xmltodict): `parse()` turns a document into dicts, and `unparse()` turns dicts back into XML. It covers the common 90% of cases and doesn't keep every XML detail, like attribute order. For exact fidelity, its README says to use lxml.
|
||||||
|
|
||||||
|
MarkupSafe [escapes characters so text is safe to use in HTML and XML](https://markupsafe.palletsprojects.com/en/latest/). `escape()` returns a `Markup` string, and any text you join to it gets escaped too. To build HTML, [use `Markup` as the format string](https://markupsafe.palletsprojects.com/en/latest/formatting/): the values you format into it are escaped first. Passing text to `Markup()` itself [marks it safe without escaping](https://markupsafe.palletsprojects.com/en/latest/escaping/#markupsafe.Markup), so save that for markup you trust.
|
||||||
|
|
||||||
|
Escape user text with MarkupSafe, and sanitize user HTML you want to keep with JustHTML. For XML you didn't write, lxml's FAQ says to [keep network access and external DTDs off](https://lxml.de/FAQ.html#how-do-i-use-lxml-safely-as-a-web-service-endpoint) and parse with `resolve_entities=False`, then reject any document that still holds entity references. With xmltodict, pass `disable_entities=True`.
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
Offering one Requests-style API for sync and async code, HTTPX spares you a second Python HTTP client. Requests suits sync scripts, aiohttp busy asyncio apps.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Sync and async code from one API, or HTTP/2: HTTPX
|
||||||
|
- Sync code with the simplest API: Requests
|
||||||
|
- An asyncio app with many concurrent requests, WebSockets, or its own HTTP server: aiohttp
|
||||||
|
- Connection pools and retries in your own hands, one layer below Requests: urllib3
|
||||||
|
- HTTPX's API, verifying TLS with your OS's certificates: HTTPX2
|
||||||
|
- Building and editing URLs: yarl
|
||||||
|
|
||||||
|
HTTPX has a [broadly Requests-compatible API](https://www.python-httpx.org/#features), sync by default with async when you need it, and runs on [asyncio or Trio](https://www.python-httpx.org/async/#supported-async-environments). Coming from Requests, [`httpx.Client` takes the place of `requests.Session`](https://www.python-httpx.org/compatibility/#client-instances). HTTP/2 is opt-in: install `httpx[http2]` and pass [`http2=True`](https://www.python-httpx.org/http2/#enabling-http2) to the client, which pays off most when you send lots of concurrent async requests.
|
||||||
|
|
||||||
|
Python's own docs [recommend Requests](https://docs.python.org/3/library/urllib.request.html) for a higher-level HTTP client. It's [sync only](https://requests.readthedocs.io/en/latest/user/advanced/#blocking-or-non-blocking), and its docs point to HTTPX, among others, for async.
|
||||||
|
|
||||||
|
aiohttp is an async client and server for asyncio, with [WebSockets on both sides](https://docs.aiohttp.org/en/stable/#key-features) and middleware for the client. Its API is wordier than Requests' by design, to [make the most of non-blocking operations](https://docs.aiohttp.org/en/stable/http_request_lifecycle.html#why-is-aiohttp-client-api-that-way). Each `async with` or `await` gives the event loop a chance to switch to other work. Install `aiohttp[speedups]` to get [aiodns for faster DNS resolving](https://docs.aiohttp.org/en/stable/#library-installation), which its docs highly recommend.
|
||||||
|
|
||||||
|
urllib3 is what gives Requests its [connection pooling](https://requests.readthedocs.io/en/latest/). Use it directly when you want the pool and the retry policy in your own code. It can [retry idempotent requests on its own](https://urllib3.readthedocs.io/en/stable/user-guide.html#retrying-requests): set the policy once on the `PoolManager` to cover every request.
|
||||||
|
|
||||||
|
HTTPX2 has [the same public API as HTTPX](https://pydantic.dev/docs/httpx2/get-started/migration/#in-a-hurry) under a new name, so switching means renaming the dependency and the import. It [verifies TLS with your operating system's trust store](https://pydantic.dev/docs/httpx2/get-started/migration/#behavior-differences) instead of a bundled certificate list. The two packages [install side by side](https://pydantic.dev/docs/httpx2/get-started/migration/#you-can-have-both-installed), but their objects don't mix: when a library takes a client, [build it from the package that library uses](https://pydantic.dev/docs/httpx2/get-started/migration/#but-objects-dont-cross-the-boundary).
|
||||||
|
|
||||||
|
yarl's `URL` is [immutable](https://yarl.aio-libs.org/en/latest/#introduction): every change returns a new URL, and strings you pass in get percent-encoded for you. Build paths with `/` and queries with `%`. Its docs pick immutability so you can [hand a URL to other code](https://yarl.aio-libs.org/en/latest/#comparison-with-other-url-libraries) without it being changed under you. aiohttp's request methods [take a yarl `URL`](https://docs.aiohttp.org/en/stable/client_quickstart.html#make-a-request) as well as a string.
|
||||||
|
|
||||||
|
Whichever client you pick, create one session object and reuse it, since it holds the connection pool: Requests' `Session`, HTTPX's `Client` or `AsyncClient`, aiohttp's `ClientSession`, urllib3's `PoolManager`. [Don't create one per request](https://docs.aiohttp.org/en/stable/client_quickstart.html#make-a-request); make one per application and pass it around. Keep the shortcut functions for [one-off scripts](https://www.python-httpx.org/advanced/clients/#why-use-a-client): HTTPX's top-level API opens a new connection for every request, and `urllib3.request()` shares one global pool with your dependencies.
|
||||||
|
|
||||||
|
Set timeouts, too. Requests [never times out unless you pass `timeout`](https://requests.readthedocs.io/en/latest/user/quickstart/#timeouts), and its docs say nearly all production code should. urllib3 [takes one on the `PoolManager`](https://urllib3.readthedocs.io/en/stable/user-guide.html#using-timeouts) to cover every request, and HTTPX [enforces timeouts by default](https://www.python-httpx.org/advanced/timeouts/).
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
Pillow handles everyday edits, so it's the default Python image processing library. Use scikit-image for scientific analysis, and pyvips when memory is tight.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Resizing, cropping, and converting images: Pillow
|
||||||
|
- Scientific image analysis on NumPy arrays: scikit-image
|
||||||
|
- Large images, or many of them, with little memory: pyvips
|
||||||
|
- ImageMagick's features from Python: Wand
|
||||||
|
- Removing photo backgrounds: rembg
|
||||||
|
- Resizing and cropping images on demand over HTTP: thumbor
|
||||||
|
- QR codes: qrcode
|
||||||
|
- Barcodes: python-barcode
|
||||||
|
|
||||||
|
Pillow is [ideal for batch processing](https://pillow.readthedocs.io/en/stable/handbook/overview.html), like making thumbnails and converting between file formats. Open each file in a `with Image.open(path) as im:` block. Opening is [fast and independent of the file size](https://pillow.readthedocs.io/en/stable/handbook/tutorial.html), since Pillow reads the pixels only when it has to. To fit an image to a size, use `ImageOps.contain()`, `cover()`, `fit()`, or `pad()`. `thumbnail()` works too, but it [changes the image in place](https://pillow.readthedocs.io/en/stable/handbook/tutorial.html#relative-resizing).
|
||||||
|
|
||||||
|
scikit-image [aims to be the reference library for scientific image analysis](https://scikit-image.org/docs/stable/about/values.html) in Python, and puts science ahead of photo editing. Images are plain NumPy arrays, so [standard NumPy operations](https://scikit-image.org/docs/stable/user_guide/numpy_images.html) work on them. To change an image's dtype, use `img_as_float()` or `img_as_ubyte()`, [never `astype`](https://scikit-image.org/docs/stable/user_guide/data_types.html), which doesn't rescale the values to the new dtype's range.
|
||||||
|
|
||||||
|
pyvips builds a pipeline of operations and runs it only when you write the result. It streams the image a section at a time, so it [doesn't keep whole images in memory](https://github.com/libvips/pyvips). To shrink an image, use `pyvips.Image.thumbnail()` instead of resizing: it [loads and resizes in one step](https://www.libvips.org/API/current/developer-checklist.html), which is faster and uses less memory. When you read an image from top to bottom, open it with [`access="sequential"`](https://libvips.github.io/pyvips/intro.html).
|
||||||
|
|
||||||
|
Wand is a [ctypes-based ImageMagick binding](https://docs.wand-py.org/en/latest/), so install ImageMagick's MagickWand library first. Its objects are resources like open files: [use them in a `with` block](https://docs.wand-py.org/en/latest/guide/resource.html) so they get closed. Wand's docs say to [never use Wand directly in an HTTP service](https://docs.wand-py.org/en/latest/guide/security.html) or on any public server. Hand the images to a background worker through a queue, and limit ImageMagick's resources and formats in its `policy.xml`.
|
||||||
|
|
||||||
|
rembg runs as a [CLI, a Python library, an HTTP server, or a Docker container](https://github.com/danielgatis/rembg). In code, create a session once with `new_session()` and pass it to each `remove()` call, since `remove` otherwise [starts a new session every call](https://github.com/danielgatis/rembg/blob/main/USAGE.md). The model weights [carry their own licenses](https://github.com/danielgatis/rembg), separate from rembg's MIT license, so check the one you use before you ship it in a commercial product.
|
||||||
|
|
||||||
|
thumbor is an HTTP server: you [set the size and crop in the image URL](https://github.com/thumbor/thumbor), and it [detects faces and important features](https://thumbor.readthedocs.io/en/latest/) to crop around them. Set a `SECURITY_KEY` so [every URL is signed](https://thumbor.readthedocs.io/en/latest/security.html) and nobody can tamper with it, and build those URLs in Python with [libthumbor](https://thumbor.readthedocs.io/en/latest/libraries.html). In production, [turn off `ALLOW_UNSAFE_URL`](https://thumbor.readthedocs.io/en/latest/configuration.html) and run [more than one instance](https://thumbor.readthedocs.io/en/latest/hosting.html) behind a load balancer.
|
||||||
|
|
||||||
|
qrcode makes a QR code in one call: `qrcode.make("Some data")`. For more control, use the `QRCode` class, and pass `version=None` with `make(fit=True)` to pick the size for you. For SVG output, the README [recommends the path factory](https://github.com/lincolnloop/python-qrcode), `SvgPathImage`. If you style the code or embed an image, set error correction to high, since styled codes aren't guaranteed to work with all readers.
|
||||||
|
|
||||||
|
python-barcode writes SVG with [no external dependencies](https://python-barcode.readthedocs.io/en/latest/). For PNG and other images, install the `python-barcode[images]` extra, which adds Pillow. Its docs [recommend SVG](https://python-barcode.readthedocs.io/en/latest/getting-started.html) unless your target can't use it, since vectors scale better. It calculates the checksum for you.
|
||||||
|
|
||||||
|
Treat every uploaded image as untrusted. Pillow [guards against decompression bombs](https://pillow.readthedocs.io/en/stable/reference/Image.html) with a pixel limit, and its `Image.open(formats=...)` argument restricts the formats it tries. With pyvips, check the image dimensions before you process it, and [block the loaders](https://www.libvips.org/API/current/developer-checklist.html) libvips hasn't tested for security. With Wand, check each file's [magic bytes](https://docs.wand-py.org/en/latest/guide/security.html), never its extension or MIME type.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
Unless you need MicroPython on a microcontroller or Pyodide in the browser, CPython is the Python implementation to run. Cython compiles your slow code to C.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Most projects: CPython
|
||||||
|
- Microcontrollers and other constrained devices: MicroPython
|
||||||
|
- Python in the browser or Node.js: Pyodide
|
||||||
|
- Slow code to compile, or a C library to wrap: Cython
|
||||||
|
- Long-running programs in pure Python: PyPy
|
||||||
|
|
||||||
|
CPython is the [original and most-maintained implementation](https://docs.python.org/3/reference/introduction.html#alternate-implementations) of Python, written in C, and new language features generally show up there first. Give each app its own [virtual environment](https://docs.python.org/3/installing/index.html) from venv, and install packages into it with pip.
|
||||||
|
|
||||||
|
MicroPython is a lean implementation of Python 3, [optimized to run on microcontrollers](https://micropython.org/) and in constrained environments. It includes a small subset of the standard library, plus modules like `machine` for the hardware. Flash your board with firmware from [the download page](https://micropython.org/download/), then manage it from your computer with [mpremote](https://docs.micropython.org/en/latest/reference/mpremote.html). `mpremote mip install` [installs packages](https://docs.micropython.org/en/latest/reference/packages.html#installing-packages-with-mpremote) from micropython-lib by default, not PyPI. When RAM runs short, [freeze the modules that rarely change](https://docs.micropython.org/en/latest/reference/packages.html#freezing-packages) into the firmware.
|
||||||
|
|
||||||
|
Pyodide is [a port of CPython to WebAssembly](https://pyodide.org/en/stable/) that runs in the browser and Node.js. It installs any pure Python wheel from PyPI, and many packages with C extensions, like NumPy and pandas, have been ported for it. Install packages with [`micropip.install()`](https://pyodide.org/en/stable/usage/loading-packages.html#how-to-chose-between-micropip-install-and-pyodide-loadpackage), which its docs recommend for almost everything. In a web page, run Pyodide [in a web worker](https://pyodide.org/en/stable/usage/webworker.html), so long computations don't freeze your UI.
|
||||||
|
|
||||||
|
Cython isn't a separate interpreter: it [translates Python code to C](https://cython.readthedocs.io/en/latest/src/quickstart/overview.html) that runs inside CPython, to speed up slow code or wrap C libraries. Write your code in [pure Python syntax](https://cython.readthedocs.io/en/latest/src/tutorial/pure.html) to keep the file runnable by the plain interpreter, and use `.pyx` files for what that syntax can't express. The annotation report from `cython -a` shows [where types help](https://cython.readthedocs.io/en/latest/src/quickstart/cythonize.html#determining-where-to-add-types). Build your package with [a build backend](https://cython.readthedocs.io/en/latest/src/userguide/source_files_and_compilation.html#compiling-with-a-build-backend), and ship [prebuilt wheels](https://cython.readthedocs.io/en/latest/src/userguide/source_files_and_compilation.html#compiling-with-pyximport) to your users.
|
||||||
|
|
||||||
|
PyPy is a replacement for CPython, and speed is the reason to use it. It works best on [long-running programs that spend much of their time in Python code](https://pypy.org/features.html#speed), not on short scripts. Code built on C extension modules is a poor fit, since they [often run much slower on PyPy than on CPython](https://doc.pypy.org/faq.html#do-c-extension-modules-work-with-pypy). Packages installed for CPython aren't available to PyPy, so [install them for PyPy](https://doc.pypy.org/faq.html#module-xyz-does-not-work-with-pypy-importerror) in its own virtual environment with `pypy -m pip`.
|
||||||
|
|
||||||
|
Before you switch implementations or compile anything for speed, [measure first](https://pypy.org/performance.html#profiling-vmprof) to confirm the slowdown is real. Then profile to find the slow parts, and only optimize those.
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
At a terminal, IPython replaces the built-in Python interactive interpreter; in a notebook, it runs in Jupyter Notebook, unless you pick marimo for .py files.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A better shell in the terminal: IPython
|
||||||
|
- Notebooks in the browser, in Python or other languages: Jupyter Notebook
|
||||||
|
- Reactive notebooks saved as `.py` files, runnable as scripts or apps: marimo
|
||||||
|
- A terminal REPL that flags syntax errors before running: ptpython
|
||||||
|
- A shell inside your own program: IPython or ptpython
|
||||||
|
|
||||||
|
IPython aims to be [an interactive shell superior to Python's default](https://ipython.readthedocs.io/en/latest/overview.html#enhanced-interactive-python-shell), with tab completion, object introspection, system shell access, and command history across sessions. Install it with [`pip install ipython`](https://ipython.readthedocs.io/en/latest/install/install.html) and start it with `ipython`. Type [`object?`](https://ipython.readthedocs.io/en/latest/interactive/tutorial.html#the-four-most-helpful-commands) for details about an object, and `object??` for more. [`%run`](https://ipython.readthedocs.io/en/latest/interactive/tutorial.html#running-and-editing) runs a script and loads its variables into your session, rereading the file each time. After an exception, [`%debug`](https://ipython.readthedocs.io/en/latest/interactive/tutorial.html#debugging) drops you into pdb to look around.
|
||||||
|
|
||||||
|
Jupyter Notebook is [a simplified notebook authoring application](https://jupyter-notebook.readthedocs.io/en/latest/). A notebook is a shareable document that mixes code, text, charts, and interactive controls. Install it with [`pip install notebook`](https://jupyter.org/install) and run `jupyter notebook`, which [opens it in your web browser](https://jupyter-notebook.readthedocs.io/en/latest/notebook.html#starting-the-notebook-server). Work on a problem [in pieces](https://jupyter-notebook.readthedocs.io/en/latest/notebook.html#basic-workflow), organizing related ideas into cells and moving on once earlier ones work. Code runs in a kernel, and [kernels for other languages](https://jupyter-notebook.readthedocs.io/en/latest/notebook.html#installing-kernels) let you write notebooks in more than Python. Notebooks save as [JSON files with the `.ipynb` extension](https://jupyter-notebook.readthedocs.io/en/latest/notebook.html#notebook-documents).
|
||||||
|
|
||||||
|
marimo is a [reactive Python notebook](https://docs.marimo.io/): run a cell, and marimo runs the cells that depend on it, keeping code and outputs consistent. Delete a cell and marimo [removes its variables from memory](https://docs.marimo.io/faq/#how-is-marimo-different-from-jupyter), so no hidden state is left behind. Each notebook is a pure Python file you can version with Git or run as a script with `python your_notebook.py`. Serve it as an app with `marimo run`, which hides the code. Install it with `pip install marimo`, then [create or edit a notebook](https://docs.marimo.io/getting_started/quickstart/) with `marimo edit your_notebook.py`. To bring a Jupyter notebook over, run `marimo convert your_notebook.ipynb -o your_notebook.py`. To know what order to run cells in, marimo [lets you define each variable in only one cell](https://docs.marimo.io/guides/coming_from/jupyter/#redefining-variables).
|
||||||
|
|
||||||
|
ptpython is [an advanced Python REPL](https://github.com/prompt-toolkit/ptpython) with syntax highlighting, multiline editing, autocompletion, and both Vi and Emacs key bindings. Before it runs your input, it [checks that it's valid Python](https://github.com/prompt-toolkit/ptpython#syntax-validation) and moves the cursor to the error. Install it with `pip install ptpython` and start it with `ptpython`. With IPython installed, [`ptipython`](https://github.com/prompt-toolkit/ptpython#ipython-support) adds IPython's magic functions and shell integration.
|
||||||
|
|
||||||
|
Both terminal shells can open inside your own program, with access to its variables: call [`IPython.embed()`](https://ipython.readthedocs.io/en/latest/interactive/reference.html#embedding-ipython) where you want to look around, or ptpython's [`embed(globals(), locals())`](https://github.com/prompt-toolkit/ptpython#embedding-the-repl). IPython is also [the default Python kernel](https://docs.jupyter.org/en/latest/what_is_jupyter.html#ipython) in Jupyter Notebook.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
A function that runs every hour fits in APScheduler inside your app. A pipeline of tasks outgrows a Python scheduler and needs Airflow, Prefect, or Dagster.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Jobs inside a running app, on cron or one-off triggers, kept across restarts: APScheduler
|
||||||
|
- Batch pipelines with a clear start and end that run on a schedule: Airflow
|
||||||
|
- Your Python functions as pipelines, with tasks created at runtime: Prefect
|
||||||
|
- Pipelines built around the data assets they produce, like tables and models: Dagster
|
||||||
|
- A simple loop of periodic jobs in one script: schedule
|
||||||
|
|
||||||
|
APScheduler [schedules your Python code to run later](https://github.com/agronholm/apscheduler/blob/3.x/README.rst), once or periodically. It runs inside your existing application, not as a service. Give each job [a trigger](https://apscheduler.readthedocs.io/en/latest/userguide.html#choosing-the-right-scheduler-job-store-s-executor-s-and-trigger-s): `date` runs it once, `interval` at fixed intervals, and `cron` at set times of day. Jobs [live in memory by default](https://apscheduler.readthedocs.io/en/latest/userguide.html#basic-concepts). When they must survive restarts and crashes, add a persistent job store.
|
||||||
|
|
||||||
|
schedule is an [in-process scheduler for periodic jobs](https://schedule.readthedocs.io/en/latest/), with no extra processes and no dependencies. Write `schedule.every(10).minutes.do(job)`, then call `schedule.run_pending()` in a loop. Its docs call it [a simple solution for simple scheduling problems](https://schedule.readthedocs.io/en/latest/#when-not-to-use-schedule), and say to look elsewhere when jobs must persist between restarts or run concurrently.
|
||||||
|
|
||||||
|
Airflow is [a platform for orchestrating batch workflows](https://airflow.apache.org/docs/apache-airflow/stable/index.html#why-airflow). Its docs say workflows with a clear start and end that run on a schedule are a great fit. It comes with a wide range of built-in operators for integrating other technologies. It's a set of services: [a minimal install](https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/overview.html#required-components) runs a scheduler, a processor that parses your workflow files, and an API server with the UI. It also needs a metadata database, usually PostgreSQL or MySQL. Write tasks with [the TaskFlow API](https://airflow.apache.org/docs/apache-airflow/stable/tutorial/taskflow.html): decorate plain Python functions, and Airflow creates the tasks, wires their dependencies, and passes data between them.
|
||||||
|
|
||||||
|
Prefect [turns your Python functions into data pipelines](https://docs.prefect.io/v3/get-started), with no DSLs or complex config files. Put [`@flow` on your script's entrypoint and `@task` on each function it calls](https://docs.prefect.io/v3/get-started/quickstart). Prefect tracks each task's state, so a failed run can resume from its point of failure. With the open-source server running, call `.serve()` on your flow with a `cron` schedule: it starts a process that runs the flow on that schedule. Its docs call serving [simple to reason about](https://docs.prefect.io/v3/concepts/deployments#static-infrastructure) for flows on a machine you control.
|
||||||
|
|
||||||
|
Dagster is [a data orchestrator built for data engineers](https://docs.dagster.io/), with lineage and observability built in. You [declare data assets like tables, datasets, and ML models as Python functions](https://github.com/dagster-io/dagster). Dagster runs them at the right time to keep them up to date. If you're just starting out, its docs [strongly recommend assets rather than ops](https://docs.dagster.io/guides/build/ops). Run assets on [a cron schedule](https://docs.dagster.io/guides/automate/schedules).
|
||||||
|
|
||||||
|
In an orchestrator, make every task safe to run twice. Airflow [can retry a failed task](https://airflow.apache.org/docs/apache-airflow/stable/best-practices.html#creating-a-task), so its docs say a task should produce the same outcome on every re-run. Prefect tasks are [retryable units of work](https://docs.prefect.io/v3/concepts/tasks) too.
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
Keep the built-in logging module until you need what a Python logging library adds: structlog's key-value logs in JSON, or Loguru's logger, ready on import.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A library other people import: logging, with only a `NullHandler`
|
||||||
|
- Key-value events, pretty in development and JSON in production: structlog
|
||||||
|
- An app already on logging that needs structured output: structlog, which wraps it
|
||||||
|
- A script you want logging in with no setup: Loguru
|
||||||
|
- Log files that rotate, expire, and compress: Loguru
|
||||||
|
|
||||||
|
Every Python module [can take part in the standard library's logging](https://docs.python.org/3/library/logging.html), so your app's log holds messages from third-party packages next to your own. In each module, create a logger with [`logging.getLogger(__name__)`](https://docs.python.org/3/howto/logging.html#advanced-logging-tutorial), so logger names follow your package layout. Use `basicConfig()` for a quick setup, and move to [`dictConfig()`](https://docs.python.org/3/howto/logging.html#configuring-logging), which the docs recommend for new applications. In a library, [add no handler other than `NullHandler`](https://docs.python.org/3/howto/logging.html#configuring-logging-for-a-library) and don't log to the root logger: handlers are for the app developer to pick.
|
||||||
|
|
||||||
|
structlog logs [events in a context of key-value pairs](https://www.structlog.org/en/latest/why.html#structured-logging) instead of prose, so each entry is a dictionary instead of a string. [Bind values to a logger](https://www.structlog.org/en/latest/getting-started.html#building-a-context), and every entry it logs carries them. In a web app, [call `clear_contextvars()` at the start of each request](https://www.structlog.org/en/latest/contextvars.html), then `bind_contextvars()` for values like a request ID. Render [pretty, colored output during development and JSON in production](https://www.structlog.org/en/latest/logging-best-practices.html#pretty-printing-vs-structured-output), since log aggregators parse JSON more easily. structlog can also [wrap the standard library's logging and add structure to it](https://www.structlog.org/en/latest/getting-started.html#structlog-and-standard-librarys-logging), so an app already built on logging keeps it.
|
||||||
|
|
||||||
|
Loguru has [one logger, ready to use](https://loguru.readthedocs.io/en/latest/overview.html#ready-to-use-out-of-the-box-without-boilerplate) after `from loguru import logger`, and it writes to stderr out of the box. Instead of handlers, formatters, and filters, you configure it with [one function, `add()`](https://loguru.readthedocs.io/en/latest/overview.html#no-handler-no-formatter-no-filter-one-function-to-rule-them-all). In an app, [call `remove()` first](https://loguru.readthedocs.io/en/latest/resources/troubleshooting.html#how-do-i-create-and-configure-a-logger) to drop the default handler, then `add()` where your logs should go. Given a file path, `add()` handles [rotation, retention, and compression](https://loguru.readthedocs.io/en/latest/overview.html#easier-file-logging-with-rotation-retention-compression), and [`serialize=True`](https://loguru.readthedocs.io/en/latest/overview.html#structured-logging-as-needed) turns each message into JSON. In a library, [call `disable()` instead of `add()`](https://loguru.readthedocs.io/en/latest/overview.html#suitable-for-scripts-and-libraries), and the app using it can `enable()` your logs again.
|
||||||
|
|
||||||
|
Whichever you pick, configure logging once, where your app starts, and have each module only get its logger. With the standard library, usually [only the root logger needs configuring](https://docs.python.org/3/library/logging.html), since module loggers pass their messages up to it. structlog wants [`structlog.configure()` on app initialization](https://www.structlog.org/en/latest/configuration.html), and with Loguru, [your other modules inherit the configuration](https://loguru.readthedocs.io/en/latest/resources/troubleshooting.html#how-do-i-create-and-configure-a-logger) from your entry point. Third-party packages that use the standard library's logging keep logging there. Send their messages to your output too: structlog [formats them with `ProcessorFormatter`](https://www.structlog.org/en/latest/standard-library.html#processor-formatter), and Loguru [intercepts them with a handler](https://loguru.readthedocs.io/en/latest/overview.html#entirely-compatible-with-standard-logging).
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
Tabular data goes to scikit-learn, and boosted trees to LightGBM or CatBoost. Both fit in that Python machine learning library's Pipeline.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Classification, regression, and clustering on tabular data: scikit-learn
|
||||||
|
- Boosted trees that train fast on large datasets: LightGBM
|
||||||
|
- Boosted trees on data with many categorical columns: CatBoost
|
||||||
|
- Bayesian networks and causal models: pgmpy
|
||||||
|
- Feature engineering on pandas dataframes: Feature-engine
|
||||||
|
- Boosted trees on a Spark, Dask, or Ray cluster: XGBoost
|
||||||
|
- Forecasting time series without training a model first: TimesFM
|
||||||
|
|
||||||
|
scikit-learn covers [supervised and unsupervised learning](https://scikit-learn.org/stable/getting_started.html), plus preprocessing, model selection, and evaluation. Deep learning is [out of its scope](https://scikit-learn.org/stable/faq.html#why-is-there-no-support-for-deep-or-reinforcement-learning-will-there-be-such-support-in-the-future). To pick a model, follow its [Choosing the right estimator](https://scikit-learn.org/stable/machine_learning_map.html) chart. Split your data into train and test sets [before any preprocessing](https://scikit-learn.org/stable/common_pitfalls.html#how-to-avoid-data-leakage). Then put the preprocessing and the model in one Pipeline, so cross-validation and search never fit on the data they score. For gradient boosting without another dependency, its HistGradientBoostingClassifier and HistGradientBoostingRegressor [handle missing values and categorical data](https://scikit-learn.org/stable/modules/ensemble.html#gradient-boosted-trees) with no preprocessing.
|
||||||
|
|
||||||
|
LightGBM aims at [faster training and lower memory use](https://lightgbm.readthedocs.io/en/latest/). It grows trees leaf-wise, so `num_leaves` is [the main parameter to tune](https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html#tune-parameters-for-the-leaf-wise-best-first-tree): keep it below 2^(max_depth). To prevent overfitting, raise `min_data_in_leaf`. Instead of one-hot encoding, mark categorical columns with `categorical_feature`: it [often performs better](https://lightgbm.readthedocs.io/en/latest/Advanced-Topics.html#categorical-feature-support). With a validation set, use [early stopping](https://lightgbm.readthedocs.io/en/latest/Python-Intro.html#early-stopping) to find the number of boosting rounds.
|
||||||
|
|
||||||
|
CatBoost takes [non-numeric features without preprocessing](https://catboost.ai/) and gives good results with its default parameters. Its docs say [not to one-hot encode](https://catboost.ai/docs/en/features/categorical-features) during preprocessing: list the categorical columns in `cat_features` instead. Before tuning anything else, [rule out underfitting and overfitting](https://catboost.ai/docs/en/concepts/parameter-tuning): set a large number of iterations, and turn on the overfitting detector and the use-best-model option. The learning rate is set from your data by default.
|
||||||
|
|
||||||
|
pgmpy does [causal and probabilistic reasoning with graphical models](https://pgmpy.org/), from learning a graph from data to running inference on the fitted model. scikit-learn [leaves graphical models out](https://scikit-learn.org/stable/faq.html#will-you-add-graphical-models-or-sequence-prediction-to-scikit-learn), so use pgmpy for Bayesian networks. Every discovery algorithm [follows one pattern](https://pgmpy.org/guides/causal_discovery.html#api): instantiate, fit, and read the result. Switching algorithms means changing only the class. Pass what you already know about the domain as [required or forbidden edges](https://pgmpy.org/guides/causal_discovery.html#expert-knowledge) with `ExpertKnowledge`. For queries, Variable Elimination is [the default choice](https://pgmpy.org/guides/probabilistic_inference.html#exact-inference) while the model is small enough for exact inference.
|
||||||
|
|
||||||
|
Feature-engine is for work where [pandas and scikit-learn are your main tools](https://feature-engine.trainindata.com/en/latest/#sitting-at-the-interface-of-pandas-and-scikit-learn). Each transformer takes the columns it changes in its `variables` argument, so it [applies steps to selected groups of variables](https://scikit-learn.org/stable/related_projects.html). Fit the transformers on the training set and transform both sets, as the [quick start](https://feature-engine.trainindata.com/en/latest/quickstart/index.html) does. Put them in a scikit-learn Pipeline, and your [whole feature engineering pipeline](https://feature-engine.trainindata.com/en/latest/quickstart/index.html#feature-engine-within-scikit-learn-s-pipeline) saves as one object.
|
||||||
|
|
||||||
|
XGBoost is built to be [efficient, flexible, and portable](https://xgboost.readthedocs.io/en/stable/). The same code runs on distributed environments, and its docs cover training on Dask, Spark, and Ray. Use its scikit-learn interface, like `XGBClassifier`, so it [works with scikit-learn's tools](https://xgboost.readthedocs.io/en/stable/python/sklearn_estimator.html) such as cross-validation. For categorical columns, pass a dataframe with the `category` dtype and set [`enable_categorical`](https://xgboost.readthedocs.io/en/stable/tutorials/categorical.html#training-with-scikit-learn-interface). Its docs suggest you tune with cross-validation, then [retrain with the best parameters and early stopping](https://xgboost.readthedocs.io/en/stable/python/sklearn_estimator.html#early-stopping). To keep a model, [save it with `save_model`](https://xgboost.readthedocs.io/en/stable/tutorials/saving_model.html), since a pickle is a memory snapshot meant only for checkpoints.
|
||||||
|
|
||||||
|
TimesFM is a forecasting model that Google Research pretrained on a large time-series corpus. It [does well zero-shot](https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/) on benchmarks from many domains, so you can forecast without training a model first. Install it with the extra for your backend, and load a checkpoint from the Hugging Face Hub. The code is Apache licensed, but [the pretrained weights carry their own license](https://github.com/google-research/timesfm), so check the model card before commercial or production use.
|
||||||
|
|
||||||
|
Every pick but TimesFM works with scikit-learn. XGBoost, LightGBM, and CatBoost ship scikit-learn estimators, Feature-engine's transformers go in a Pipeline, and pgmpy is [scikit-learn compatible where possible](https://pgmpy.org/). Independent benchmarks find no single winner among the three boosting libraries, so compare them on your own data in the same cross-validation.
|
||||||
|
|
||||||
|
Treat a model file you didn't make like code. scikit-learn's docs say to [never load a pickle from an untrusted source](https://scikit-learn.org/stable/model_persistence.html#security-maintainability-limitations), and point to skops.io or ONNX instead. XGBoost's [security notes](https://xgboost.readthedocs.io/en/stable/security.html#use-of-python-pickle) say the same about pickles.
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
Instead of a Python messaging library per broker, an asyncio service can run on FastStream. Otherwise, use confluent-kafka, pika, or paho-mqtt.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- An asyncio service on Kafka, RabbitMQ, MQTT, NATS, or Redis: FastStream
|
||||||
|
- Kafka, including Confluent's Schema Registry: confluent-kafka
|
||||||
|
- RabbitMQ, with the client its team recommends: pika
|
||||||
|
- IoT devices and sensors that speak MQTT: paho-mqtt
|
||||||
|
|
||||||
|
confluent-kafka is Confluent's Kafka client, [a binding on top of librdkafka](https://docs.confluent.io/kafka-clients/python/current/overview.html), the C client. `produce()` only queues a message, so [call `poll()` to get delivery reports](https://docs.confluent.io/kafka-clients/python/current/overview.html#asynchronous-writes). Its README says production apps [should serialize with Schema Registry](https://github.com/confluentinc/confluent-kafka-python#basic-producer-example) instead of producing raw bytes.
|
||||||
|
|
||||||
|
pika is [a pure-Python implementation of AMQP 0-9-1](https://pika.readthedocs.io/en/stable/), and RabbitMQ's tutorials use it as [the Python client the RabbitMQ team recommends](https://www.rabbitmq.com/tutorials/tutorial-one-python). Start with `BlockingConnection`, as those tutorials do: it's the adapter for [a procedural style](https://pika.readthedocs.io/en/stable/modules/adapters/index.html). [Ack each message after processing it](https://www.rabbitmq.com/tutorials/tutorial-two-python#message-acknowledgment), so RabbitMQ redelivers it when a worker dies.
|
||||||
|
|
||||||
|
paho-mqtt is the Eclipse Foundation's MQTT client, and MQTT is [a lightweight publish/subscribe protocol](https://eclipse.dev/paho/files/paho.mqtt.python/html/index.html) for IoT devices and low-bandwidth links. Create a client, connect it, and run [one of its network loops](https://eclipse.dev/paho/files/paho.mqtt.python/html/index.html#network-loop). `loop_forever()` blocks until you disconnect, and `loop_start()` runs the loop in a background thread. Both reconnect for you. [Subscribe in `on_connect()`](https://eclipse.dev/paho/files/paho.mqtt.python/html/index.html#getting-started), so a reconnect renews your subscriptions.
|
||||||
|
|
||||||
|
FastStream is [an asynchronous framework for event-driven services](https://faststream.ag2.ai/latest/): FastAPI's decorators, type-driven validation, and dependency injection, pointed at Kafka, RabbitMQ, NATS, Redis, or MQTT instead of HTTP. Write handlers with `@broker.subscriber()` and `@broker.publisher()`, run the app with `faststream run`, and [test it in memory](https://faststream.ag2.ai/latest/#testing-the-service) with no broker running. It [wraps your broker's own client](https://faststream.ag2.ai/latest/#your-broker-in-full) and leaves out retries, delayed delivery, and task orchestration by design: for background jobs, see [Task Queues](/categories/task-queues/).
|
||||||
|
|
||||||
|
In asyncio code, pick a client that won't block the event loop. FastStream is asynchronous throughout, confluent-kafka has [AsyncIO clients](https://docs.confluent.io/kafka-clients/python/current/overview.html#when-to-use-asyncio-clients) for apps like a FastAPI service, and pika has [an asyncio connection adapter](https://pika.readthedocs.io/en/stable/modules/adapters/asyncio.html).
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
On Windows, pywin32 gives Python the Win32 API and COM, from registry keys to Excel. Calling .NET needs a Python Windows library of its own: pythonnet.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Win32 API calls, the registry, the event log, or COM automation of apps like Excel: pywin32
|
||||||
|
- .NET assemblies from Python, or Python embedded in a .NET app: pythonnet
|
||||||
|
- Several Python versions on one machine, switched per folder: pyenv-win
|
||||||
|
- A portable Python that runs from a folder, a network share, or a USB key: WinPython
|
||||||
|
|
||||||
|
pywin32 provides [access to many of the Windows APIs, including COM](https://github.com/mhammond/pywin32). Install it with `python -m pip install --upgrade pywin32`. Outside a virtual environment, its post-install script sets up COM objects and services. [Never run it inside one](https://github.com/mhammond/pywin32#installing-via-pip). To drive an app over COM, [pass its name to `win32com.client.Dispatch()`](https://mhammond.github.io/pywin32/html/com/win32com/HTML/QuickStartClientCom.html), like `"Excel.Application"`. Then call its methods and set its properties as on any Python object.
|
||||||
|
|
||||||
|
pythonnet [integrates CPython with the .NET runtime](https://pythonnet.github.io/pythonnet/) instead of compiling Python to .NET code, so your existing Python code and C extensions keep working. Install it with `pip install pythonnet`. [Pick the runtime before you import `clr`](https://pythonnet.github.io/pythonnet/python.html#loading-a-runtime), with `load("coreclr")` or the `PYTHONNET_RUNTIME` environment variable, or you get the default one. Then [load each assembly with `clr.AddReference()`](https://pythonnet.github.io/pythonnet/python.html#importing-modules) and import its namespaces like Python packages. It also works the other way, to [embed Python in a .NET app](https://pythonnet.github.io/pythonnet/dotnet.html).
|
||||||
|
|
||||||
|
pyenv-win [brings pyenv to Windows](https://github.com/pyenv-win/pyenv-win#introduction), so you get pyenv's commands for switching between Python versions. Install it with [the PowerShell script from its quick start](https://github.com/pyenv-win/pyenv-win#quick-start), then set your default with `pyenv install <version>` and `pyenv global <version>`. In a project folder, [`pyenv local <version>`](https://github.com/pyenv-win/pyenv-win#usage) picks the Python that `python` runs there, with nothing to activate.
|
||||||
|
|
||||||
|
WinPython is a [portable Python distribution](https://winpython.github.io/). Unzip it anywhere, like a network share or a USB key, and it runs, with nothing installed or written to the registry. Pick a build by [how much comes preinstalled](https://winpython.github.io/#builds): a bare Python you fill yourself, or the scientific stack with Spyder and JupyterLab ready to open. [Add packages with pip or wppm](https://winpython.github.io/#about), then zip the folder and hand it to your students or colleagues.
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
Some helpers aren't in the standard library, so before writing your own Python utility library, check boltons. Blinker dispatches signals between modules.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Helpers the standard library lacks, like atomic file saves and chunked iteration: boltons
|
||||||
|
- Events inside one process, between modules that don't import each other: Blinker
|
||||||
|
|
||||||
|
boltons is a set of [pure-Python utilities missing from the standard library](https://boltons.readthedocs.io/en/latest/), like atomic file saves and chunked iteration. Run `pip install boltons`, then [import what you need from its modules](https://boltons.readthedocs.io/en/latest/#installation-and-integration), e.g. `from boltons.cacheutils import LRU`. It depends on no other packages and its modules are self-contained, so you can also [copy the whole package into your project, or just one module](https://boltons.readthedocs.io/en/latest/architecture.html#integration). Most modules aim to be good enough for basic uses. When you outgrow one, [its docs often point to a third-party alternative](https://boltons.readthedocs.io/en/latest/#third-party-packages).
|
||||||
|
|
||||||
|
Blinker lets [any number of interested parties subscribe to events, or signals](https://github.com/pallets-eco/blinker), inside one Python process. Create a named signal with [`signal('name')`](https://blinker.readthedocs.io/en/latest/#decoupling-with-named-signals): every call with that name returns the same object, so modules and plugins share it without importing each other. Register receivers with `connect()`, and when you call `send()`, [pass the object that emits the signal](https://blinker.readthedocs.io/en/latest/#emitting-signals), usually `self`. A receiver can then [subscribe to one sender only](https://blinker.readthedocs.io/en/latest/#subscribing-to-specific-senders).
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
Whether you're building an app or learning the field, there's a Python NLP library for it: spaCy for apps, NLTK for learning. Stanza covers many languages.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Production NLP pipelines: spaCy
|
||||||
|
- Learning NLP, or lexical resources like WordNet: NLTK
|
||||||
|
- Many languages, or CoreNLP from Python: Stanza
|
||||||
|
- Chinese word segmentation: jieba
|
||||||
|
- Chinese characters to pinyin: pypinyin
|
||||||
|
- Spaces between CJK text and letters or digits: pangu.py
|
||||||
|
|
||||||
|
spaCy is [designed for production use](https://spacy.io/usage/facts-figures#comparison-usage): you build and train NLP pipelines, then package them to deploy. Its trained pipelines are Python packages. In an automated build, [install them with pip from a direct link](https://spacy.io/usage/models#download-pip) instead of spaCy's download command, and put that link in your requirements.txt. In a larger code base, [import the pipeline as a module](https://spacy.io/usage/models#models-loading), so a missing one raises an ImportError right away. Run your texts through `nlp.pipe` [in batches](https://spacy.io/usage/processing-pipelines#processing), and disable the components you don't need. To train your own pipeline, run [`spacy train` with one config file](https://spacy.io/usage/training#quickstart).
|
||||||
|
|
||||||
|
NLTK comes with [corpora and lexical resources such as WordNet](https://www.nltk.org/), plus libraries to tokenize, tag, parse, and classify text. Its creators wrote a book that teaches NLP with it, and you can read it online. The book also says NLTK [isn't highly optimized for runtime performance](https://www.nltk.org/book/ch00.html#natural-language-toolkit-nltk), so use it to learn and experiment. The data ships separately: [get it with NLTK's data downloader](https://www.nltk.org/data.html), and set `NLTK_DATA` when you install it somewhere other than the standard locations.
|
||||||
|
|
||||||
|
Stanza is [designed to work across many languages, using the Universal Dependencies formalism](https://stanfordnlp.github.io/stanza/#about). It's also the official Python interface to Stanford's Java CoreNLP, though its own pipeline doesn't need CoreNLP. You can [list the processors to load](https://stanfordnlp.github.io/stanza/getting_started.html#specifying-processors) with `processors=`. Pass all your documents to the pipeline [at once](https://stanfordnlp.github.io/stanza/getting_started.html#processing-multiple-documents), since a for loop over one sentence at a time is very slow. For a lot of text, [run it on a GPU](https://stanfordnlp.github.io/stanza/getting_started.html#controlling-devices). To keep the pipeline from downloading anything at runtime, [download the models ahead of time](https://stanfordnlp.github.io/stanza/getting_started.html#downloading-models-for-offline-usage).
|
||||||
|
|
||||||
|
jieba [segments Chinese text into words](https://github.com/fxsjy/jieba): accurate mode suits text analysis, and search engine mode cuts long words into short ones for better recall. Add your own words with `jieba.load_userdict()` to get higher accuracy, and for Traditional Chinese, switch to its bigger dictionary with `jieba.set_dictionary()`. The other picks work with it: spaCy can [use jieba as its Chinese segmenter](https://spacy.io/usage/models#chinese), and Stanza [supports it as a tokenizer](https://stanfordnlp.github.io/stanza/pipeline.html).
|
||||||
|
|
||||||
|
pypinyin [matches pinyin by whole words](https://github.com/mozillazg/python-pinyin), so it handles characters with more than one reading. It also writes zhuyin (Bopomofo) and Wade-Giles. When a wrong word split gives a wrong reading, [segment the text with jieba first](https://pypinyin.readthedocs.io/zh-cn/latest/faq.html) and pass in the list of words. For readings that are still wrong, [add your own](https://pypinyin.readthedocs.io/zh-cn/latest/usage.html#custom-dict) with `load_phrases_dict()` or `load_single_dict()`.
|
||||||
|
|
||||||
|
pangu.py [inserts spaces between CJK characters and letters, digits, and symbols](https://github.com/vinta/pangu.py). Call `pangu.space_text()` on a string or `pangu.space_file()` on a file. From the command line, `pangu-py -c` prints the corrected text and exits with 1 when the spacing needed fixing.
|
||||||
|
|
||||||
|
Check the license of the models and data, not only the library. spaCy is [MIT](https://github.com/explosion/spaCy/blob/master/LICENSE), but its [Spanish pipelines](https://spacy.io/models/es) are GPL and its [Italian ones](https://spacy.io/models/it) are for non-commercial use only. NLTK is [Apache](https://github.com/nltk/nltk/blob/develop/LICENSE.txt), and its corpora come [under various licenses](https://github.com/nltk/nltk/wiki/FAQ), each listed in its own README.
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
Configure network devices from several vendors through one API with the Python network automation library NAPALM. Scapy forges, sends, and sniffs packets.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Configuring network devices from several vendors, or reading their state: NAPALM
|
||||||
|
- Forging, sending, sniffing, and dissecting network packets: Scapy
|
||||||
|
|
||||||
|
NAPALM [talks to network devices of different operating systems through a unified API](https://napalm.readthedocs.io/en/latest/), to change their configuration or retrieve data from them. `pip install napalm` [installs all the core drivers](https://napalm.readthedocs.io/en/latest/installation/index.html). Pick the driver for your device's OS with [`get_network_driver()`](https://napalm.readthedocs.io/en/latest/#selecting-the-right-driver), like `get_network_driver('eos')`. Then [connect in a `with` block](https://napalm.readthedocs.io/en/latest/tutorials/context_manager.html), which opens and closes the session for you. To change the configuration, [load a candidate that replaces the device's config or merges into it](https://napalm.readthedocs.io/en/latest/tutorials/changing_the_config.html#replacing-the-configuration), and check the diff with `compare_config()`. Then apply it with `commit_config()`, or throw it away with `discard_config()`. What each driver supports differs, so [test your workflow in a lab first](https://napalm.readthedocs.io/en/latest/support/index.html#configuration-support-matrix).
|
||||||
|
|
||||||
|
Scapy [sends, sniffs, dissects, and forges network packets](https://scapy.readthedocs.io/en/latest/introduction.html#about-scapy), and you can use it [as a shell or as a library](https://github.com/secdev/scapy). Instead of telling you a port is open or closed, it [gives you the full decoded packets](https://scapy.readthedocs.io/en/latest/introduction.html#what-makes-scapy-so-special), so you probe once and interpret many times. Install it with `pip install scapy`, and on Windows, [install Npcap first](https://scapy.readthedocs.io/en/latest/installation.html#windows). [Sending packets needs root privileges](https://scapy.readthedocs.io/en/latest/usage.html#starting-scapy), so start the shell with `sudo scapy`, or on Windows from a command prompt with administrator privileges. Build a packet by [stacking layers with `/`](https://scapy.readthedocs.io/en/latest/usage.html#stacking-layers), like `IP()/TCP()`, and fields you don't set get sensible defaults. In your own scripts, [import what you need from `scapy.all`](https://scapy.readthedocs.io/en/latest/extending.html#using-scapy-in-your-tools).
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
SQLAlchemy is the Python ORM for most projects, and Django projects use the Django ORM. On MongoDB, pick Beanie for async code and MongoEngine for sync.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A Django project: Django ORM
|
||||||
|
- Any other app on a relational database: SQLAlchemy
|
||||||
|
- A FastAPI app with simple tables: SQLModel
|
||||||
|
- A small app that wants one module and no dependencies: peewee
|
||||||
|
- MongoDB from async code: Beanie
|
||||||
|
- MongoDB from sync code: MongoEngine
|
||||||
|
- Amazon DynamoDB: PynamoDB
|
||||||
|
|
||||||
|
SQLAlchemy follows the [data mapper pattern](https://www.sqlalchemy.org/philosophy.html), so your classes and your schema can change separately. Open the Session [in a `with` block](https://docs.sqlalchemy.org/en/latest/orm/quickstart.html), and [keep its lifecycle outside](https://docs.sqlalchemy.org/en/latest/orm/session_basics.html) the functions that read or write data: one Session per thread, one AsyncSession per task. To avoid N+1 queries, load related objects up front with [eager loading](https://docs.sqlalchemy.org/en/latest/orm/queryguide/relationships.html).
|
||||||
|
|
||||||
|
The Django ORM follows the [Active Record pattern](https://docs.djangoproject.com/en/stable/misc/design-philosophies/): each model holds both the fields and the behavior of its data. Before you optimize, [find out what queries you run](https://docs.djangoproject.com/en/stable/topics/db/optimization/) and what they cost.
|
||||||
|
|
||||||
|
SQLModel is [designed for FastAPI apps](https://sqlmodel.tiangolo.com/): one class is both a Pydantic model and a SQLAlchemy model. Get the session from a [FastAPI dependency](https://sqlmodel.tiangolo.com/tutorial/fastapi/session-with-dependency/), so each request gets its own. When you need more complex features, [plug SQLAlchemy in directly](https://sqlmodel.tiangolo.com/features/).
|
||||||
|
|
||||||
|
For a small app, peewee is a [single module with no required dependencies](https://docs.peewee-orm.com/en/latest/), with few concepts to learn. In a web app, [open a connection per request](https://docs.peewee-orm.com/en/latest/peewee/database.html) and close it when the request ends. Once traffic grows, [switch to a pooled database](https://docs.peewee-orm.com/en/latest/peewee/framework_integration.html).
|
||||||
|
|
||||||
|
On the NoSQL side, Beanie is an [async MongoDB ODM built on Pydantic](https://beanie-odm.dev/), with one Document class per collection. Pass your document models to [`init_beanie()`](https://beanie-odm.dev/tutorial/initialization/), which creates the collections and the indexes you declared.
|
||||||
|
|
||||||
|
MongoEngine is the sync pick: it's [built on PyMongo only](https://mongoengine-odm.readthedocs.io/faq.html) and doesn't support async drivers. Its document schemas are [enforced in your app, not by MongoDB](https://mongoengine-odm.readthedocs.io/tutorial.html), and they catch wrong types and missing fields.
|
||||||
|
|
||||||
|
PynamoDB is a [Pythonic interface to DynamoDB](https://pynamodb.readthedocs.io/en/stable/). Build on its [Model API](https://pynamodb.readthedocs.io/en/stable/tutorial.html): one model class per table, each with a hash key. For tests, point it at a [local DynamoDB-compatible server](https://pynamodb.readthedocs.io/en/stable/local.html).
|
||||||
|
|
||||||
|
Track schema changes in migrations: Django [has them built in](https://docs.djangoproject.com/en/stable/topics/migrations/), SQLAlchemy has [Alembic](https://alembic.sqlalchemy.org/en/latest/), and Beanie [ships its own](https://beanie-odm.dev/tutorial/migrations/). Review every migration Alembic's autogenerate writes, since it's [not meant to be perfect](https://alembic.sqlalchemy.org/en/latest/autogenerate.html).
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
Whatever you need from PyPI, uv installs and locks it, and uv-build packages pure Python. Conda is more than a Python package manager: it handles any language.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A project's dependencies, lockfile, and Python versions: uv
|
||||||
|
- Packaging a pure-Python project: uv-build
|
||||||
|
- Non-Python dependencies, like C libraries or R: conda
|
||||||
|
- Nothing beyond what ships with Python: pip
|
||||||
|
- A project already on Poetry: Poetry
|
||||||
|
- Tests across several Python versions: Hatch
|
||||||
|
- Command-line apps, each in its own environment: pipx
|
||||||
|
- C or C++ extension modules: setuptools
|
||||||
|
- Build hooks or a flexible project layout: Hatchling
|
||||||
|
|
||||||
|
uv is one tool for [your project's dependencies, lockfile, and Python versions](https://docs.astral.sh/uv/), and it covers what you'd otherwise use pip, pipx, and Poetry for. Start with `uv init`, add dependencies with `uv add`, and run your code with `uv run`, which [syncs the environment with the lockfile](https://docs.astral.sh/uv/guides/projects/#running-commands) first. Its `uv pip` commands are meant [for projects not ready to move away from pip](https://docs.astral.sh/uv/pip/).
|
||||||
|
|
||||||
|
uv-build is uv's own build backend, which its docs call [a great choice for most Python projects](https://docs.astral.sh/uv/concepts/build-backend/#choosing-a-build-backend). It's for pure-Python code, aims to need no configuration, and by default expects your package in `src/<package_name>/`.
|
||||||
|
|
||||||
|
conda manages [packages in any language](https://packaging.python.org/en/latest/key_projects/#conda), Python included, so it fits projects that need compiled non-Python libraries next to their Python ones. Install it with [Miniforge](https://docs.conda.io/projects/conda/en/stable/), which uses the free conda-forge channel. Miniconda and the Anaconda Distribution use Anaconda's repository instead, which [may require a commercial license](https://docs.conda.io/projects/conda/en/stable/user-guide/getting-started.html#before-you-start). In a conda environment, [install as much as you can with conda](https://docs.conda.io/projects/conda/en/stable/user-guide/tasks/manage-environments.html#using-pip-in-an-environment) before you use pip for the rest. To change the environment later, create a new one rather than running conda after pip.
|
||||||
|
|
||||||
|
pip is the [standard tool to install packages from PyPI](https://packaging.python.org/en/latest/guides/tool-recommendations/#installing-packages), and it ships with most Python installations. Run it as `python -m pip`, so [it installs into the interpreter you name](https://pip.pypa.io/en/stable/user_guide/#running-pip), and use it inside a [virtual environment](https://pip.pypa.io/en/stable/getting-started/#next-steps). To repeat an install, [pin every version](https://pip.pypa.io/en/stable/topics/repeatable-installs/#pinning-the-package-versions) in a requirements file generated by `pip freeze`. By default, pip doesn't check downloads for tampering, so for deployments its docs [recommend hash-checking mode](https://pip.pypa.io/en/stable/topics/secure-installs/).
|
||||||
|
|
||||||
|
Poetry [manages dependencies and packaging](https://python-poetry.org/docs/), with a lockfile for repeatable installs. [Install it in its own virtual environment](https://python-poetry.org/docs/#installation), with pipx for example, and never in the project it manages.
|
||||||
|
|
||||||
|
Hatch is a [Python project manager](https://hatch.pypa.io/latest/) for environments, builds, and publishing. You never create its environments by hand: the first command you run in one [creates it](https://hatch.pypa.io/latest/environment/). Define a [matrix](https://hatch.pypa.io/latest/config/environment/advanced/#matrix) to run the same tests across Python versions. For most projects, `hatch test` [runs pytest and coverage.py](https://hatch.pypa.io/latest/tutorials/testing/overview/) with no test environment of your own.
|
||||||
|
|
||||||
|
pipx installs command-line apps [each in its own virtual environment](https://pipx.pypa.io/stable/#pip-vs-pipx) and puts their commands on your PATH, and `pipx run` runs an app without installing it. If you already use uv, its tool command does the same job. pipx's docs suggest pipx [when you need its extras](https://pipx.pypa.io/stable/explanation/comparisons.html#picking-one), like installing an app system-wide.
|
||||||
|
|
||||||
|
setuptools [builds C and C++ extension modules](https://setuptools.pypa.io/en/latest/userguide/ext_modules.html), and Hatch's docs [recommend it](https://hatch.pypa.io/latest/why/#build-backend) when you need them. Keep your config in `pyproject.toml` and [only the dynamic parts](https://setuptools.pypa.io/en/latest/userguide/quickstart.html#setuppy-discouraged) in `setup.py`. Build with `python -m build` [instead of running `setup.py`](https://packaging.python.org/en/latest/discussions/setup-py-deprecated/).
|
||||||
|
|
||||||
|
Hatchling is the build backend the Python Packaging User Guide's [tutorial uses by default](https://packaging.python.org/en/latest/tutorials/packaging-projects/#choosing-a-build-backend). It supports plugins and [build hooks](https://hatch.pypa.io/latest/config/build/#build-hooks), and uv's docs point to it when you need build scripts or a more flexible project layout.
|
||||||
|
|
||||||
|
Whichever tool builds your package, declare the backend in `pyproject.toml`'s `[build-system]` table and your metadata in [the standard `[project]` table](https://packaging.python.org/en/latest/guides/writing-pyproject-toml/), which most build backends understand. For an app, commit [uv.lock](https://docs.astral.sh/uv/guides/projects/#uvlock) or [poetry.lock](https://python-poetry.org/docs/basic-usage/#as-an-application-developer), so every machine installs the same versions. A library's lockfile doesn't reach the apps that install it, since they [resolve its dependencies themselves](https://python-poetry.org/docs/basic-usage/#as-a-library-developer).
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
Your own Python package repository can mirror PyPI with bandersnatch, or host private packages over a PyPI cache with devpi. Warehouse is the code behind PyPI.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A full or filtered mirror of PyPI: bandersnatch
|
||||||
|
- Private packages, with PyPI served from the same index: devpi
|
||||||
|
- A PyPI cache that keeps working offline: devpi
|
||||||
|
- Testing and staging your releases before PyPI: devpi
|
||||||
|
- Learning how PyPI works, or contributing to it: Warehouse
|
||||||
|
|
||||||
|
bandersnatch is what [PyPI's help page recommends](https://pypi.org/help/#mirroring) for running your own mirror. It only syncs the static files installers need: it takes no uploads and skips PyPI's dynamic APIs. [Its quickstart](https://github.com/pypa/bandersnatch#quickstart) has you run `bandersnatch mirror` once to write a config file, then adapt that file. Run it again to fill the mirror, and on a schedule to keep it current. [Add an allowlist or blocklist](https://bandersnatch.readthedocs.io/en/latest/filtering_configuration.html#allowlist-blocklist-filtering-settings) to cut the mirror down, and serve its `web/` directory with any static web server.
|
||||||
|
|
||||||
|
devpi is what [PyPI's help page recommends](https://pypi.org/help/#private-indices) for private packages, which PyPI doesn't host. devpi-server sets up `root/pypi`, [a caching mirror of PyPI](https://devpi.net/docs/devpi/devpi/stable/+doc/quickstart-pypimirror.html) that downloads each release on first request and works offline after that. When a cache is all you need, point pip at it. For your own packages, [create an index with `root/pypi` as its base](https://devpi.net/docs/devpi/devpi/stable/+doc/quickstart-releaseprocess.html#initializing-a-basic-server-and-index), so one URL serves your uploads and all of PyPI. By default, a name you upload there [hides the PyPI package with the same name](https://devpi.net/docs/devpi/devpi/stable/+doc/userman/devpi_indices.html#modifying-the-mirror-whitelist), which stops dependency confusion attacks. devpi-client handles the release steps: `devpi upload`, `devpi test` to run their tox tests, and `devpi push` to a staging index or on to PyPI. Before you put the server on the internet, [secure it](https://devpi.net/docs/devpi/devpi/stable/+doc/adminman/security.html): its docs warn that exposing it isn't safe by default.
|
||||||
|
|
||||||
|
Warehouse powers PyPI itself, and [its own docs say](https://warehouse.pypa.io/application/#usage-assumptions-and-concepts) people who run their own package index usually use other tools, like devpi. Read it to learn how PyPI works, or [set up its development environment](https://warehouse.pypa.io/development/getting-started/) with Docker to contribute.
|
||||||
|
|
||||||
|
Whichever you run, point pip at your server as its only index. [pip's docs call `--extra-index-url` unsafe](https://pip.pypa.io/en/stable/cli/pip_install/#cmdoption-extra-index-url) for private packages: pip checks every index with no priority, so a package with the same name on PyPI can win. And [serve your repository over valid HTTPS](https://packaging.python.org/en/latest/guides/hosting-your-own-index/).
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
Only run sqlmap and SET against targets you're authorized to test. Like mitmproxy and Sherlock, they're Python penetration testing tools with one job each.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Testing a web app for SQL injection: sqlmap
|
||||||
|
- Testing user awareness in social-engineering assessments: Social-Engineer Toolkit (SET)
|
||||||
|
- Inspecting, changing, or scripting HTTP and HTTPS traffic: mitmproxy
|
||||||
|
- Finding which social networks a username is registered on: Sherlock
|
||||||
|
|
||||||
|
sqlmap [automates detecting and exploiting SQL injection flaws](https://github.com/sqlmapproject/sqlmap) and taking over database servers. Its docs prefer [cloning the Git repository](https://github.com/sqlmapproject/sqlmap/wiki/Download-and-update) to installing from PyPI: `git clone --depth 1 https://github.com/sqlmapproject/sqlmap.git sqlmap-dev`. Run it from that checkout, where `python sqlmap.py -h` lists the basic options, and read the [user's manual](https://github.com/sqlmapproject/sqlmap/wiki/Usage) for the rest.
|
||||||
|
|
||||||
|
The Social-Engineer Toolkit (SET) is [a penetration testing framework for authorized social-engineering assessments](https://github.com/trustedsec/social-engineer-toolkit). It gives security teams guided attack vectors to test user awareness and run consent-based red-team exercises. On Kali Linux or WSL, install it with `sudo apt install set`. Elsewhere, install a source checkout into a virtual environment. Then `sudo setoolkit` launches its interactive console.
|
||||||
|
|
||||||
|
mitmproxy is [an interactive, TLS-capable intercepting proxy](https://docs.mitmproxy.org/stable/). It comes as three front ends to one core: the mitmproxy console, the mitmweb browser GUI, and mitmdump on the command line. The first two keep every flow in memory, so they're for small samples, while mitmdump records traffic and transforms it programmatically. On macOS, [install it](https://docs.mitmproxy.org/stable/overview/installation/) with `brew install --cask mitmproxy`. On Linux and Windows, download it from mitmproxy.org. It [starts as a regular HTTP proxy on localhost:8080](https://docs.mitmproxy.org/stable/overview/getting-started/): point your browser or device at it, then browse to mitm.it and install mitmproxy's certificate authority to see HTTPS traffic too. To change traffic in code, write a Python [addon](https://docs.mitmproxy.org/stable/addons/overview/#anatomy-of-an-addon) and load it with `-s`.
|
||||||
|
|
||||||
|
Sherlock [hunts down social media accounts by username](https://sherlockproject.xyz/) across social networks. Its docs [suggest pipx over pip](https://sherlockproject.xyz/installation): `pipx install sherlock-project`. Then `sherlock user123` [searches for one username](https://sherlockproject.xyz/usage), and `sherlock user1 user2 user3` for several. It saves the accounts it finds to a text file named after each username.
|
||||||
|
|
||||||
|
sqlmap and SET both put permission first. sqlmap calls [attacking targets without prior mutual consent illegal](https://github.com/sqlmapproject/sqlmap/wiki/License), and its FAQ says to practice [only against a target you own or have explicit permission to test](https://github.com/sqlmapproject/sqlmap/wiki/FAQ#where-can-i-practise-using-sqlmap). SET is [only for authorized testing](https://github.com/trustedsec/social-engineer-toolkit#responsible-use) where explicit permission and scope have been established, never against systems, accounts, networks, or people without consent.
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
Qiskit for most circuits, PennyLane for circuits you train, Cirq for device-level control, QuTiP for physics: Python quantum computing libraries you can mix.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Most circuit work, including IBM hardware: Qiskit
|
||||||
|
- Circuits you train with gradients, like quantum machine learning: PennyLane
|
||||||
|
- Device-level control over circuits and noise: Cirq
|
||||||
|
- Simulating the dynamics of open quantum systems: QuTiP
|
||||||
|
|
||||||
|
Qiskit is an [SDK for working with quantum computers](https://github.com/Qiskit/qiskit) at the level of circuits, operators, and primitives. Its docs lay out the development workflow as a [Qiskit pattern](https://quantum.cloud.ibm.com/docs/en/guides/intro-to-patterns): map your problem to circuits, optimize them for the target hardware, run them, and post-process the results. Run circuits through the [Sampler and Estimator primitives](https://quantum.cloud.ibm.com/docs/en/guides/primitives). Sampler samples outcomes, and Estimator estimates expectation values. Before a circuit goes to a device, [transpile it](https://quantum.cloud.ibm.com/docs/en/guides/transpile) to the gates and qubit connections that device supports. Test it in [local testing mode](https://quantum.cloud.ibm.com/docs/en/guides/local-testing-mode) first: once your program works there, moving to a QPU takes only a backend name change.
|
||||||
|
|
||||||
|
PennyLane lets you [train a quantum computer the same way as a neural network](https://docs.pennylane.ai/en/stable/): it differentiates quantum circuits and connects them to PyTorch, JAX, and NumPy. Write each circuit as a Python function under the [`qnode` decorator](https://docs.pennylane.ai/en/stable/introduction/circuits.html#the-qnode-decorator), which ties it to the device that runs it. Pick the [interface](https://docs.pennylane.ai/en/stable/introduction/interfaces.html) for your machine learning library, then train with that library's own optimizers.
|
||||||
|
|
||||||
|
Cirq is for [writing, manipulating, and optimizing quantum circuits](https://quantumai.google/cirq) when the details of the hardware matter. A circuit is [a collection of Moments](https://quantumai.google/cirq/build/circuits), each a set of operations that act in the same time slice. Model your target processor as a [Device](https://quantumai.google/cirq/hardware/devices) and validate your circuits against it. Then [compile them with transformers](https://quantumai.google/cirq/transform/transformers) into circuits that device can run. Test small circuits on the [built-in simulators](https://quantumai.google/cirq/simulate/simulation), then move to qsim, which its docs recommend for most users. Before Google hardware, run on the [Quantum Virtual Machine](https://quantumai.google/cirq/simulate/quantum_virtual_machine), which mimics Google's processors with noise data.
|
||||||
|
|
||||||
|
QuTiP [simulates the dynamics of open quantum systems](https://qutip.org/), like the ones in quantum optics, trapped ions, and superconducting circuits. Every state and operator is a [`Qobj`](https://qutip.readthedocs.io/en/stable/guide/guide-basics.html), and QuTiP has predefined ones for a variety of them. Pick the solver [by the kind of system](https://qutip.readthedocs.io/en/stable/guide/dynamics/dynamics-intro.html): a closed system is a state vector, and an open one needs a density matrix. [`mesolve`](https://qutip.readthedocs.io/en/stable/guide/dynamics/dynamics-master.html) covers both, and switches to the master equation when you give it collapse operators. For large systems, the docs recommend the [Monte Carlo solver](https://qutip.readthedocs.io/en/stable/guide/dynamics/dynamics-monte.html). To simulate quantum circuits, use its [qutip-qip](https://qutip.readthedocs.io/en/stable/guide/guide-family.html) package.
|
||||||
|
|
||||||
|
You don't have to stay with one library. PennyLane [imports Qiskit circuits](https://docs.pennylane.ai/en/stable/introduction/circuits.html#importing-circuits-from-other-frameworks), and its plugins run on [Qiskit devices](https://docs.pennylane.ai/projects/qiskit/en/stable/) and [Cirq's simulators](https://docs.pennylane.ai/projects/cirq/en/stable/). QuTiP's qutip-qip [simulates circuits made in Qiskit](https://qutip-qip.readthedocs.io/en/stable/qip-qiskit.html).
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
Match the Python recommender system library to your data: Surprise if users rate items, implicit if they only buy or watch. Annoy finds similar items fast.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Ratings users give, like 1 to 5 stars: Surprise
|
||||||
|
- Purchases, watches, or page views, with no ratings: implicit
|
||||||
|
- Fast lookups of similar items from a trained model: Annoy
|
||||||
|
|
||||||
|
Surprise is built for [explicit rating data](https://surpriselib.com/), and it doesn't support implicit ratings. Its [basic usage](https://surprise.readthedocs.io/en/latest/getting_started.html#basic-usage) cross-validates an algorithm like SVD on a built-in dataset in a few lines of code. For your own ratings, define a `Reader` and [load them from a file or a pandas dataframe](https://surprise.readthedocs.io/en/latest/getting_started.html#use-a-custom-dataset). Surprise predicts ratings, so to recommend items, [predict the ratings a user hasn't given and keep the top N](https://surprise.readthedocs.io/en/latest/FAQ.html#how-to-get-the-top-n-recommendations-for-each-user).
|
||||||
|
|
||||||
|
implicit offers fast Python implementations of popular algorithms for [implicit feedback datasets](https://benfred.github.io/implicit/), such as Alternating Least Squares and Bayesian Personalized Ranking. Every model [implements one interface](https://benfred.github.io/implicit/api/models/index.html) for training and recommending. Train with `fit` on a [sparse CSR matrix of users by items](https://benfred.github.io/implicit/api/models/recommender_base.html#implicit.recommender_base.RecommenderBase.fit), whose values are how confident you are that the user likes each item. Then `recommend` returns items for a user, and `similar_items` finds related items.
|
||||||
|
|
||||||
|
Annoy is a C++ library with Python bindings that [searches for points close to a query point](https://github.com/spotify/annoy). For recommendations, it runs [after matrix factorization](https://github.com/spotify/annoy#background): every user and item becomes a vector, and Annoy finds similar ones. Its indexes are static files, so you build an index once and save it, and every process [loads it with mmap](https://github.com/spotify/annoy#python-code-example) and shares the same data.
|
||||||
|
|
||||||
|
implicit and Annoy work together: implicit can use an Annoy index to [speed up `recommend` and `similar_items`](https://benfred.github.io/implicit/api/ann.html) on any matrix factorization model, at the risk of missing some relevant results.
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
Underneath most Python scientific computing libraries sits NumPy's array. SciPy adds algorithms on top, and Numba compiles loops that won't vectorize.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Arrays and vectorized math: NumPy
|
||||||
|
- Optimization, integration, interpolation, and linear algebra: SciPy
|
||||||
|
- Numeric loops that won't vectorize: Numba
|
||||||
|
- Exact, symbolic math: SymPy
|
||||||
|
- Regression, time series, and hypothesis tests: statsmodels
|
||||||
|
- Biological sequences and file formats: Biopython; molecules: RDKit
|
||||||
|
- Physical units: Pint; astronomy: Astropy; seismology: ObsPy
|
||||||
|
- Bayesian models: PyMC
|
||||||
|
- Discrete-event simulations: SimPy; agent-based models: Mesa
|
||||||
|
- Graphs and networks: NetworkX
|
||||||
|
- Planar geometry: Shapely
|
||||||
|
- Colour science: Colour; math animations: Manim
|
||||||
|
|
||||||
|
NumPy gives you a multidimensional array and fast routines that run on it. Write whole-array expressions instead of Python loops: [vectorization and broadcasting](https://numpy.org/doc/stable/user/whatisnumpy.html#why-is-numpy-fast) push the element-by-element work into compiled C code, and your code reads closer to the math. For random numbers, [create a Generator with `default_rng()`](https://numpy.org/doc/stable/reference/random/index.html#random-quick-start) and call its methods.
|
||||||
|
|
||||||
|
SciPy is [a collection of algorithms built on NumPy](https://docs.scipy.org/doc/scipy/tutorial/index.html), split into subpackages by domain, like `optimize`, `integrate`, `interpolate`, `linalg`, and `stats`. Its docs recommend [using those subpackages as namespaces](https://docs.scipy.org/doc/scipy/reference/index.html#guidelines-for-importing-functions-from-scipy): `from scipy import optimize`, then `optimize.curve_fit()`. For linear algebra, use `scipy.linalg` over `numpy.linalg`. It has everything NumPy's module has and more, and it's [always compiled with BLAS/LAPACK](https://docs.scipy.org/doc/scipy/tutorial/linalg.html#scipy-linalg-vs-numpy-linalg), so it may run faster.
|
||||||
|
|
||||||
|
Numba is a just-in-time compiler that [works best on code that uses NumPy arrays and functions, and loops](https://numba.readthedocs.io/en/stable/user/5minguide.html). NumPy wants vector operations, but Numba is [happy with plain loops](https://numba.readthedocs.io/en/stable/user/performance-tips.html#loops), so it fits the numeric code you can't vectorize. Put `@jit` on the function and [let Numba decide when and how to optimize](https://numba.readthedocs.io/en/stable/user/jit.html#lazy-compilation). Keep pandas out of those functions: Numba doesn't understand it, so that code [runs in the interpreter](https://numba.readthedocs.io/en/stable/user/5minguide.html#will-numba-work-for-my-code) with Numba's overhead on top.
|
||||||
|
|
||||||
|
SymPy does symbolic math: expressions stay [exact, not approximate](https://docs.sympy.org/latest/tutorials/intro-tutorial/intro.html#what-is-symbolic-computation), and it's written entirely in Python. Its best practices keep symbolic and numeric code apart: [model the problem in SymPy, then turn the result into a function with `lambdify()`](https://docs.sympy.org/latest/explanation/best-practices.html#separate-symbolic-and-numeric-code) that runs on NumPy arrays. Define symbols with `symbols()` and [the assumptions you know](https://docs.sympy.org/latest/explanation/best-practices.html#defining-symbols), like `positive=True`, so more expressions simplify. Build expressions from those symbols, [not from strings](https://docs.sympy.org/latest/explanation/best-practices.html#avoid-string-inputs).
|
||||||
|
|
||||||
|
statsmodels [estimates statistical models, runs hypothesis tests, and explores data](https://www.statsmodels.org/stable/index.html), and [most of its results are verified](https://www.statsmodels.org/stable/about.html#testing) against another statistical package. SciPy's stats docs [send regression, linear models, and time series analysis to statsmodels](https://docs.scipy.org/doc/scipy/reference/stats.html), while `scipy.stats` keeps the probability distributions and statistical tests. Write your model as an [R-style formula](https://www.statsmodels.org/stable/example_formulas.html) on a pandas DataFrame, like `smf.ols("y ~ x", data=df).fit()`, and read the fit with `summary()`.
|
||||||
|
|
||||||
|
Biopython [parses bioinformatics file formats](https://biopython.org/docs/latest/Tutorial/chapter_introduction.html) like FASTA and GenBank, queries online services like NCBI, and gives you a standard sequence class. Read sequence files with `Bio.SeqIO.parse()`, or [`Bio.SeqIO.read()` when a file holds one record](https://biopython.org/docs/latest/Tutorial/chapter_seqio.html#parsing-or-reading-sequences). When you query NCBI through `Bio.Entrez`, [pass your email](https://biopython.org/docs/latest/Tutorial/chapter_entrez.html#entrez-guidelines) so NCBI can contact you if there's a problem. RDKit is a [cheminformatics toolkit with its core in C++](https://www.rdkit.org/docs/Overview.html), for 2D and 3D molecular operations and descriptors for machine learning. To compare molecules, [create a fingerprint generator](https://www.rdkit.org/docs/GettingStartedInPython.html#fingerprinting-and-molecular-similarity) for the fingerprint type you want, the docs' most consistent way to get fingerprints.
|
||||||
|
|
||||||
|
Pint handles [physical quantities](https://pint.readthedocs.io/en/stable/getting/overview.html): a value times a unit, in any numeric type. In a package, [create one `UnitRegistry` in a single place](https://pint.readthedocs.io/en/stable/getting/pint-in-your-projects.html#having-a-shared-registry) and import it everywhere, since quantities from different registries don't mix. Astropy is [a common core package for astronomy](https://www.astropy.org/), with affiliated packages around it. [Import the subpackage you need](https://docs.astropy.org/en/stable/importing_astropy.html#importing-astropy-and-sub-packages), like `from astropy import units as u`, and never import with `*`. ObsPy is [a framework for processing seismological data](https://docs.obspy.org/): `read()` loads [SAC, MiniSEED, and other formats into a Stream](https://docs.obspy.org/tutorial/code_snippets/reading_seismograms.html) of Traces, and its [FDSN client](https://docs.obspy.org/packages/obspy.clients.fdsn.html#basic-fdsn-client-usage) fetches data from data centers.
|
||||||
|
|
||||||
|
PyMC [builds Bayesian models with a simple Python API](https://www.pymc.io/welcome.html) and fits them with Markov chain Monte Carlo or variational inference. Declare the variables inside a `with pm.Model():` block, which [adds them to the model for you](https://www.pymc.io/projects/docs/en/stable/learn/core_notebooks/pymc_overview.html#model-specification), and let `pm.sample()` pick the samplers. Validate the model with [posterior predictive checks](https://www.pymc.io/projects/docs/en/stable/learn/core_notebooks/posterior_predictive.html), and run prior predictive checks too, which its docs call a crucial part of the Bayesian workflow.
|
||||||
|
|
||||||
|
SimPy runs discrete-event simulations. Each process is a Python generator that [yields events and waits for them](https://simpy.readthedocs.io/en/latest/simpy_intro/basic_concepts.html), and shared resources model congestion points like servers and checkout counters. Request a resource in a `with` block so it's [released for you](https://simpy.readthedocs.io/en/latest/simpy_intro/shared_resources.html#basic-resource-usage). Its docs call it [overkill for fixed-step simulations](https://simpy.readthedocs.io/en/latest/) whose processes don't interact or share resources. Mesa builds [agent-based models](https://mesa.readthedocs.io/latest/) from core components like spatial grids, shows them in the browser, and hands the results to Python's data tools. Its best practices put [the model class in `model.py` and the agents in `agents.py`](https://mesa.readthedocs.io/latest/best-practices.html#model-layout), with an optional browser visualization in `app.py`. Its [DataCollector](https://mesa.readthedocs.io/latest/getting_started.html) collects model-level and agent-level data.
|
||||||
|
|
||||||
|
NetworkX is for [the creation, manipulation, and study of complex networks](https://networkx.org/documentation/stable/index.html). [Pick the graph class first](https://networkx.org/documentation/stable/reference/introduction.html#graphs): Graph, DiGraph, MultiGraph, or MultiDiGraph. Store numeric edge data under [the `weight` key](https://networkx.org/documentation/stable/reference/introduction.html#nodes-and-edges), which algorithms like Dijkstra's shortest path read by default. Its drawing functions are basic, since its [main goal is graph analysis, not visualization](https://networkx.org/documentation/stable/reference/drawing.html), so export the graph to a dedicated visualization tool.
|
||||||
|
|
||||||
|
Shapely wraps the GEOS library for [manipulation and analysis of planar geometric objects](https://shapely.readthedocs.io/en/stable/). For many geometries, put them in a NumPy array and call its [vectorized functions instead of a Python loop](https://shapely.readthedocs.io/en/stable/#usage). It doesn't read or write data files or transform coordinate systems. Operations [presume the features sit in the same Cartesian plane](https://shapely.readthedocs.io/en/stable/manual.html), so project your data to a plane before you measure it.
|
||||||
|
|
||||||
|
Colour provides [algorithms and datasets for colour science](https://colour.readthedocs.io/en/latest/). Manim [generates animations of technical concepts from Python code](https://docs.manim.community/en/stable/). It comes in more than one edition, and the one listed here is [the community-maintained one its docs recommend](https://docs.manim.community/en/stable/faq/installation.html#which-version-should-i-use), especially for beginners. Each animation lives in [the `construct()` method of a `Scene` subclass](https://docs.manim.community/en/stable/tutorials/quickstart.html), with helper functions outside the class.
|
||||||
|
|
||||||
|
Keep your data in NumPy arrays as it moves between most of these libraries. ObsPy already holds each trace's samples in one, and Colour [recommends them as input](https://colour.readthedocs.io/en/latest/basics.html). For results you can reproduce, seed your random numbers: [pass a seed to `default_rng()`](https://numpy.org/doc/stable/reference/random/index.html#random-quick-start) in NumPy, and [give your Mesa model a `seed` argument](https://mesa.readthedocs.io/latest/best-practices.html#randomization).
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
The engine comes first, then its official Python search library. Meilisearch keeps a search bar simple, while analytics call for Elasticsearch or OpenSearch.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A typo-tolerant search bar for your site or app: meilisearch
|
||||||
|
- Log analytics and aggregations beyond search: elasticsearch
|
||||||
|
- The same under Apache instead of the AGPL, or on Amazon OpenSearch Service: opensearch-py
|
||||||
|
- Search over Django models, with a backend you can swap later: django-haystack
|
||||||
|
|
||||||
|
Meilisearch is [an open-source search engine](https://github.com/meilisearch/meilisearch-python), and meilisearch is its Python client. Its docs call it [a perfect choice for a typo-tolerant search bar](https://www.meilisearch.com/docs/resources/comparisons/alternatives), built for instant search aimed at end users. Meilisearch [queues writes and processes them asynchronously](https://www.meilisearch.com/docs/capabilities/indexing/tasks_and_batches/async_operations), so after `add_documents()`, the Python quick start [waits for indexing to complete](https://www.meilisearch.com/docs/getting_started/sdks/python) with `client.wait_for_task(task.task_uid)`.
|
||||||
|
|
||||||
|
For log analytics and aggregations beyond search, Meilisearch's own docs point you to [Elasticsearch](https://www.meilisearch.com/docs/resources/comparisons/elasticsearch#when-to-choose-elasticsearch) or [OpenSearch](https://www.meilisearch.com/docs/resources/comparisons/opensearch#when-to-choose-opensearch). elasticsearch is Elastic's official Python client. It's [unopinionated and extensible](https://www.elastic.co/docs/reference/elasticsearch/clients/python) and covers the entire Elasticsearch API. Its docs call [the bulk helpers the recommended way to ingest data](https://www.elastic.co/docs/reference/elasticsearch/clients/python/getting-started): unlike calling `client.bulk` yourself, they handle retries and send documents chunk by chunk.
|
||||||
|
|
||||||
|
OpenSearch is [a fork of Elasticsearch, and all of its software is Apache-licensed](https://opensearch.org/faq/). The Elasticsearch server is available [under the AGPL or source-available licenses](https://www.elastic.co/pricing/faq/licensing). OpenSearch's Python client is opensearch-py, [a fork of elasticsearch-py](https://github.com/opensearch-project/opensearch-py) that [wraps the OpenSearch REST API in Python methods](https://docs.opensearch.org/latest/clients/python-low-level/#low-level-python-client). Use it for any OpenSearch cluster, as OpenSearch's docs [recommend OpenSearch clients for OpenSearch clusters](https://docs.opensearch.org/latest/clients/#legacy-clients). On Amazon OpenSearch Service, [sign requests with `AWSV4SignerAuth`](https://docs.opensearch.org/latest/clients/python-low-level/#connecting-to-amazon-opensearch-service) and your IAM credentials.
|
||||||
|
|
||||||
|
django-haystack gives Django [one API over pluggable search backends](https://django-haystack.readthedocs.io/en/latest/), such as Elasticsearch, so you can switch engines without rewriting your search code. Its FAQ [advises against it](https://django-haystack.readthedocs.io/en/latest/faq.html#when-should-i-not-be-using-haystack) for data that isn't in Django models and for ultra-high volume, since the abstraction costs performance and some engine features. [Create a `SearchIndex` for each model](https://django-haystack.readthedocs.io/en/latest/tutorial.html#creating-searchindexes) you index, in a `search_indexes.py` file in its app. Load your data with [`./manage.py rebuild_index`](https://django-haystack.readthedocs.io/en/latest/tutorial.html#reindex), then keep the index current with a cron job running `update_index`.
|
||||||
|
|
||||||
|
Keep your data in your database, and send the search engine a copy of what people search for. Meilisearch [wasn't designed to be your main data container](https://www.meilisearch.com/docs/capabilities/indexing/advanced/indexing_best_practices#do-not-use-meilisearch-as-your-main-database), and django-haystack's docs call the database [authoritative and the search index non-authoritative](https://django-haystack.readthedocs.io/en/latest/signal_processors.html).
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
When your Python serialization library should validate too, msgspec decodes JSON straight into typed objects. Without a schema, orjson is fast and correct.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Decoding into typed, validated objects: msgspec
|
||||||
|
- Faster JSON that still gives you dicts and lists: orjson
|
||||||
|
- Compact binary data other languages can read: msgpack
|
||||||
|
- Schemas that convert ORM or app objects for a web API: marshmallow
|
||||||
|
|
||||||
|
msgspec works as a faster JSON or MessagePack library on its own. Still, its docs recommend it for the [full serialization and validation workflow](https://github.com/msgspec/msgspec): define your schemas with type annotations, encode them, then decode them back with validation. [Structs are the preferred way](https://msgspec.dev/structs) to define those types. Pass the type when you decode, and msgspec [validates while decoding](https://msgspec.dev/usage#typed-decoding) at no added runtime cost: `msgspec.json.decode(data, type=User)`. A consistent interface covers MessagePack, YAML, and TOML too.
|
||||||
|
|
||||||
|
Moving from the standard `json` module to orjson? The [largest difference](https://github.com/ijl/orjson?tab=readme-ov-file#migrating) is that `orjson.dumps` returns `bytes`, not `str`. orjson serializes dataclasses, datetimes, and UUIDs natively. Decoding gives you only dicts, lists, and other builtins, and orjson leaves schemas to [validation libraries a level above](https://github.com/ijl/orjson?tab=readme-ov-file#questions).
|
||||||
|
|
||||||
|
msgpack reads and writes MessagePack, a binary format that [lets you exchange data among languages like JSON, but faster and smaller](https://github.com/msgpack/msgpack-python). Call `packb` and `unpackb` for one-shot use, and read many objects from one stream with an [`Unpacker`](https://github.com/msgpack/msgpack-python?tab=readme-ov-file#streaming-unpacking). When the data comes from an untrusted source, [set `max_buffer_size`](https://msgpack-python.readthedocs.io/en/latest/api.html#msgpack.Unpacker) to limit the buffer.
|
||||||
|
|
||||||
|
marshmallow [makes no assumption about your web framework or database layer](https://marshmallow.readthedocs.io/en/latest/why.html#agnostic), so its schemas work with just about any ORM, or none. [Declare a schema](https://marshmallow.readthedocs.io/en/latest/quickstart.html#declaring-schemas) as a class that maps attribute names to fields. Its `dump` method turns your objects into primitive Python types, and `load` [validates and deserializes](https://marshmallow.readthedocs.io/en/latest/quickstart.html#deserializing-objects-loading) incoming data, raising `ValidationError` on invalid input. To get objects back instead of dicts, [decorate a schema method with `post_load`](https://marshmallow.readthedocs.io/en/latest/quickstart.html#deserializing-to-objects).
|
||||||
|
|
||||||
|
orjson, msgspec, and msgpack all pitch speed, but msgspec's own benchmark page [encourages you to write your own benchmarks](https://msgspec.dev/benchmarks) before deciding.
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
Of the two Python static site generators here, Pelican builds your blog. Nikola also builds sites that aren't only a blog, and takes Jupyter notebooks as posts.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A blog: Pelican
|
||||||
|
- A site with more than a blog, or no blog at all: Nikola
|
||||||
|
- Jupyter notebooks as posts: Nikola
|
||||||
|
- Themes written in Mako: Nikola
|
||||||
|
|
||||||
|
Pelican splits content into [articles and pages](https://docs.getpelican.com/en/latest/content.html#articles-and-pages): articles are dated, like blog posts, and pages hold content that rarely changes, like an About page. Start a project with [`pelican-quickstart`](https://docs.getpelican.com/en/latest/quickstart.html#create-a-project). Then write each post as a Markdown or reStructuredText file in `content/`, with metadata headers at the top. While you write, [`pelican --autoreload --listen`](https://docs.getpelican.com/en/latest/publish.html#site-generation) rebuilds the site on every change and serves it on localhost. Themes are Jinja2 templates. The default theme is plain HTML without styling, so start from a [community theme](https://docs.getpelican.com/en/latest/themes.html) or write your own. Add features with [plugins you install with pip](https://docs.getpelican.com/en/latest/plugins.html#how-to-use-plugins), which Pelican can discover on its own. It [isn't only for blogs](https://docs.getpelican.com/en/latest/faq.html#is-pelican-only-suitable-for-blogs), though a site without one takes some theme and configuration changes.
|
||||||
|
|
||||||
|
Nikola [started as a blog generator but supports most kinds of sites](https://getnikola.com/handbook.html#what-s-nikola-and-what-can-you-do-with-it). It comes with [batteries included](https://getnikola.com/): comments, tags, archives, feeds, multilingual support, image galleries, and code listings. Out of the box, it takes reStructuredText, Markdown, Jupyter notebooks, and HTML. Create a site with [`nikola init --demo`](https://getnikola.com/getting-started.html#init), which runs a setup wizard. Write each post with [`nikola new_post -e`](https://getnikola.com/getting-started.html#newpost), which adds the metadata headers for you. Posts default to reStructuredText; add `-f markdown` for Markdown. `nikola build` fills the `output` directory, and [`nikola serve --browser`](https://getnikola.com/getting-started.html#serve) opens the site in your browser. For a server that rebuilds on every change, run `nikola auto --browser`. Themes are [Mako or Jinja2 templates](https://getnikola.com/theming.html). For a site with no blog, follow its [guide to building a site that isn't a blog](https://getnikola.com/creating-a-site-not-a-blog-with-nikola.html).
|
||||||
|
|
||||||
|
Both write plain HTML files: Pelican's output is [easy to host anywhere](https://docs.getpelican.com/en/latest/), and a Nikola site runs [on any web server](https://getnikola.com/). For GitHub Pages, push your Pelican sources and let [a GitHub Actions workflow](https://docs.getpelican.com/en/latest/tips.html#publishing-to-github-pages-using-a-custom-github-actions-workflow) build the site. Nikola's [`nikola github_deploy`](https://getnikola.com/handbook.html#deploying-to-github) builds the site and pushes the output to a `gh-pages` branch. Moving off WordPress works with either: Pelican [imports a WordPress XML export](https://docs.getpelican.com/en/latest/importer.html), and Nikola has [`nikola import_wordpress`](https://getnikola.com/handbook.html#importing-your-wordpress-site-into-nikola).
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
Audit your dependencies for known vulnerabilities as part of Python supply chain security. A uv project has uv audit built in; other projects run pip-audit.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A project managed by uv: uv audit
|
||||||
|
- Environments, requirements files, and projects outside uv: pip-audit
|
||||||
|
|
||||||
|
`uv audit` comes with uv. Run it in your project, and it [audits the project's dependencies](https://docs.astral.sh/uv/reference/cli/#uv-audit) for known vulnerabilities and for statuses like deprecation and quarantine. By default, it covers every extra and dependency group.
|
||||||
|
|
||||||
|
pip-audit [scans Python environments for packages with known vulnerabilities](https://github.com/pypa/pip-audit), using the Python Packaging Advisory Database. Run `pip-audit` for the current environment, `pip-audit -r requirements.txt` for a requirements file, or `pip-audit .` for a local project. In CI, run it with its [official GitHub Action](https://github.com/pypa/pip-audit#github-actions). Only audit a requirements file you would install, since `pip-audit -r` is [functionally equivalent to `pip install -r`](https://github.com/pypa/pip-audit#security-model).
|
||||||
|
|
||||||
|
Also control what gets installed, since an audit only finds vulnerabilities someone has already reported. For requirements files, pip's docs recommend [hash-checking mode](https://pip.pypa.io/en/stable/topics/secure-installs/) to protect against remote tampering. In a uv project, uv's docs suggest a [dependency cooldown](https://docs.astral.sh/uv/concepts/resolution/#dependency-cooldowns), which holds back new releases until the community has had a chance to vet them.
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
Two questions pick a Python task queue: is your app async, and what's your broker? Taskiq is for asyncio. Celery takes RabbitMQ or Redis, RQ Redis, Huey SQLite.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- An asyncio app, like FastAPI: Taskiq
|
||||||
|
- A batteries-included queue on RabbitMQ, Redis, or Amazon SQS, or tasks sent from non-Python code: Celery
|
||||||
|
- The simplest job queue, on Redis or Valkey: RQ
|
||||||
|
- No broker server to run, with the queue in SQLite or the Postgres you already have: Huey
|
||||||
|
- A production backend for Django's task framework: Huey
|
||||||
|
- Retries with backoff and acks after processing by default, on RabbitMQ or Redis: Dramatiq
|
||||||
|
|
||||||
|
Taskiq's docs call it [an asyncio Celery implementation](https://taskiq-python.github.io/): you mark tasks with `@broker.task` and send them with `await task.kiq()`. Since sending is async only, its docs [point fully synchronous projects to Celery or Dramatiq](https://taskiq-python.github.io/guide/). For production, its docs [highly recommend](https://taskiq-python.github.io/guide/getting-started.html) taskiq-aio-pika or taskiq-nats as the broker and taskiq-redis as the result backend. In a FastAPI app, [taskiq-fastapi](https://taskiq-python.github.io/framework_integrations/taskiq-with-fastapi.html) connects the broker to your app.
|
||||||
|
|
||||||
|
Celery is a task queue [with batteries included](https://docs.celeryq.dev/en/stable/getting-started/first-steps-with-celery.html): [celery beat](https://docs.celeryq.dev/en/stable/userguide/periodic-tasks.html) runs periodic tasks, and [Django support](https://docs.celeryq.dev/en/stable/django/first-steps-with-django.html) comes out of the box. It runs on RabbitMQ, Redis, or Amazon SQS. Its docs call RabbitMQ [an excellent choice for production](https://docs.celeryq.dev/en/stable/getting-started/first-steps-with-celery.html#choosing-a-broker) and warn that Redis is more likely to lose data on a sudden shutdown or power failure. Its protocol has clients in other languages, so non-Python code can send tasks too. Celery [doesn't support Windows](https://docs.celeryq.dev/en/stable/getting-started/introduction.html), though.
|
||||||
|
|
||||||
|
RQ is a simple job queue on Redis or Valkey, built for [a low barrier to entry](https://python-rq.org/). [Any Python function call](https://python-rq.org/docs/) can go on a queue, and there are [no queues, exchanges, or routing rules](https://python-rq.org/docs/#on-the-design) to set up first. Jobs are pickled, so RQ is Python-only. Put job functions in a module the worker [can import, not in `__main__`](https://python-rq.org/docs/#considerations-for-jobs), and run workers on the same source code as your app. Start workers with `rq worker --with-scheduler` to run jobs you schedule with `enqueue_in` or `enqueue_at`.
|
||||||
|
|
||||||
|
Huey is [a lightweight alternative](https://huey.readthedocs.io/en/latest/) with zero dependencies, and its queue can live in Redis, Postgres, SQLite, files, or memory. Its docs [match the storage to your workload](https://huey.readthedocs.io/en/latest/guide.html#storage-options): Redis for busy workloads, SQLite for moderate ones without a separate server, and Postgres when your app already runs it. Recurring tasks come built in with the `periodic_task()` decorator. For Django, Huey has [its own integration](https://huey.readthedocs.io/en/latest/contrib.html#django), and it provides [a production backend for Django's task framework](https://huey.readthedocs.io/en/latest/contrib.html#django-task-framework).
|
||||||
|
|
||||||
|
Dramatiq aims for [sane defaults for most SaaS workloads](https://dramatiq.io/motivation.html) on RabbitMQ or Redis. Mark a function with `@dramatiq.actor` and enqueue it with `.send()`, [passing only JSON-encodable arguments](https://dramatiq.io/guide.html). When an actor raises, Dramatiq [retries it with exponential backoff](https://dramatiq.io/guide.html#error-handling). It [acknowledges a message only after processing it](https://dramatiq.io/advanced.html#message-persistence), while Celery by default does so [just before running the task](https://docs.celeryq.dev/en/stable/userguide/tasks.html#Task.acks_late). For cron-style jobs, pair it with [a separate scheduler](https://dramatiq.io/cookbook.html#scheduling). Dramatiq is [licensed under the LGPL](https://dramatiq.io/).
|
||||||
|
|
||||||
|
Make tasks safe to run twice: after a worker failure, the same message [can arrive again](https://dramatiq.io/best_practices.html#retriable-actors). Pass a task the ID of a database row rather than the object, and [re-fetch it when the task runs](https://docs.celeryq.dev/en/stable/userguide/tasks.html#state), since old data leads to race conditions. In Django, enqueue a task only [after the transaction commits](https://docs.celeryq.dev/en/stable/userguide/tasks.html#database-transactions), with Celery's `delay_on_commit()` or Huey's `on_commit_task()`. With celery beat, Huey, or Taskiq, run [only one scheduler](https://docs.celeryq.dev/en/stable/userguide/periodic-tasks.html) for periodic tasks, or you'll get duplicate tasks. In production, run workers under [a process manager](https://python-rq.org/docs/workers/).
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
If your templates should hold no Python code, Jinja fits: a Python template engine whose sandbox also renders untrusted templates. Mako embeds plain Python.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Templates without embedded Python: Jinja
|
||||||
|
- Templates your users write: Jinja, in its sandbox
|
||||||
|
- Plain Python inside your templates: Mako
|
||||||
|
|
||||||
|
Jinja [doesn't allow arbitrary Python code in templates](https://jinja.palletsprojects.com/en/stable/faq/): you get blocks, filters, and function calls, and the rest of your logic stays in Python. Outside a framework, [create one `Environment`](https://jinja.palletsprojects.com/en/stable/api/) when your app starts and load templates through a `PackageLoader`. Put the layout your pages share in a base template, and let each page [extend it and override its blocks](https://jinja.palletsprojects.com/en/stable/templates/#template-inheritance). Flask [sets up Jinja for you](https://flask.palletsprojects.com/en/stable/templating/#jinja-setup), Django has a [built-in Jinja2 backend](https://docs.djangoproject.com/en/stable/topics/templates/#django.template.backends.jinja2.Jinja2), and FastAPI [supports it with `Jinja2Templates`](https://fastapi.tiangolo.com/advanced/templates/).
|
||||||
|
|
||||||
|
When your users write templates, like a report layout, render them in Jinja's [sandbox](https://jinja.palletsprojects.com/en/stable/sandbox/#sandbox), which can block attribute access, method calls, and other operations. Its docs say the sandbox alone [is not a solution for perfect security](https://jinja.palletsprojects.com/en/stable/sandbox/#security-considerations): catch errors when a template renders, limit CPU and memory, and pass the template only the data it needs.
|
||||||
|
|
||||||
|
Mako is [an embedded Python language](https://www.makotemplates.org/): inside `<% %>` tags you [write regular Python](https://docs.makotemplates.org/en/latest/syntax.html#python-blocks). A real application [loads its templates from a `TemplateLookup`](https://docs.makotemplates.org/en/latest/usage.html#using-templatelookup) with a `module_directory`, which caches each compiled template on disk as a Python module.
|
||||||
|
|
||||||
|
Jinja leaves HTML escaping [off by default](https://jinja.palletsprojects.com/en/stable/faq/#why-is-html-escaping-not-the-default), since it also renders plain text, emails, and config files. When you create the `Environment` yourself, turn it on with [`select_autoescape()`](https://jinja.palletsprojects.com/en/stable/api/#jinja2.select_autoescape); Flask turns it on for HTML templates, and Django's Jinja2 backend turns it on for all. Mako escapes HTML only through [the `h` filter](https://docs.makotemplates.org/en/latest/filtering.html#expression-filtering): add it to `default_filters` on your `TemplateLookup` to escape every expression.
|
||||||
@@ -0,0 +1,41 @@
|
|||||||
|
Plain assert statements and fixtures make pytest the Python testing framework for new code. Add Hypothesis for property-based tests, Playwright for browsers.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Writing tests: pytest, plus Hypothesis for edge cases; Robot Framework for non-programmers
|
||||||
|
- Tests across Python versions: tox, or Nox to configure them in Python
|
||||||
|
- Browser tests: Playwright; Selenium or SeleniumBase for WebDriver suites
|
||||||
|
- Load tests written in Python: Locust
|
||||||
|
- Tests generated from an OpenAPI or GraphQL schema: Schemathesis
|
||||||
|
- Mocking: unittest.mock; responses, RESPX, or VCR.py for HTTP; FreezeGun for time
|
||||||
|
- Test objects: factory_boy for ORM models, Polyfactory for type hints
|
||||||
|
- Code coverage: Coverage.py
|
||||||
|
- Fake data: Faker or Mimesis
|
||||||
|
|
||||||
|
pytest lets you write tests with [plain `assert` statements](https://docs.pytest.org/en/stable/) and shows you what failed. It also runs your unittest suites as they are, so you can [move an old suite over bit by bit](https://docs.pytest.org/en/stable/how-to/unittest.html). Share setup through [fixtures](https://docs.pytest.org/en/stable/how-to/fixtures.html): use `yield` fixtures for teardown, and put the ones several test modules need in `conftest.py`. For a new project, pytest's docs recommend a src layout and the [importlib import mode](https://docs.pytest.org/en/stable/explanation/goodpractices.html).
|
||||||
|
|
||||||
|
Hypothesis adds property-based tests to pytest or unittest. You describe the inputs with a strategy passed to [`@given`](https://hypothesis.readthedocs.io/en/latest/quickstart.html), and Hypothesis picks which ones to try, including edge cases you didn't think of. It's [an addition to unit tests, not always a replacement](https://hypothesis.readthedocs.io/en/latest/tutorial/introduction.html): start with round trips like encode/decode, and with tests you already parametrize. Use the [most general strategy](https://hypothesis.readthedocs.io/en/latest/explanation/domain.html) your test should pass for.
|
||||||
|
|
||||||
|
Robot Framework is a [keyword-driven framework for acceptance testing](https://robotframework.org/robotframework/latest/RobotFrameworkUserGuide.html). Tests are tables of keywords, and you build higher-level keywords out of existing ones. That suits teams where people who don't write Python read or write the tests.
|
||||||
|
|
||||||
|
tox and Nox both run your tests in separate virtual environments, one per Python version or task. Configure tox [in TOML](https://tox.wiki/en/latest/tutorial/getting-started.html), in `tox.toml` or `pyproject.toml`. It tests the installed package, not your checkout, so it [catches packaging mistakes](https://docs.pytest.org/en/stable/explanation/goodpractices.html). Nox is configured in Python, in a `noxfile.py`, and tox's own docs point you to it [if tox configuration is too limiting](https://tox.wiki/en/latest/explanation.html).
|
||||||
|
|
||||||
|
Playwright was [created for end-to-end testing](https://playwright.dev/python/docs/intro) and runs Chromium, Firefox, and WebKit. Write your tests with its pytest plugin, which gives each test its own browser context. Playwright [waits for elements to be ready](https://playwright.dev/python/docs/actionability) before each action, so you don't add waits yourself. Find elements [by role, text, or test id](https://playwright.dev/python/docs/locators) rather than CSS or XPath, which break when the page changes.
|
||||||
|
|
||||||
|
Selenium drives real browsers through WebDriver, on your machine or on remote ones through Selenium Grid. Keep it for the WebDriver suites you already have. It [doesn't structure your test suite for you](https://www.selenium.dev/documentation/test_practices/), so run it under a test runner like pytest, and use [explicit waits](https://www.selenium.dev/documentation/webdriver/waits/) for the exact condition you need. SeleniumBase builds on Selenium's WebDriver APIs and [runs under pytest](https://github.com/seleniumbase/SeleniumBase), and its methods wait for elements that need time to load.
|
||||||
|
|
||||||
|
For load tests, Selenium's docs [advise against using it](https://www.selenium.dev/documentation/test_practices/discouraged/performance_testing/); use Locust. You [write the tests in regular Python code](https://docs.locust.io/en/stable/what-is-locust.html): a `User` class with `@task` methods. When you need more load, [run one worker per CPU core](https://docs.locust.io/en/stable/running-distributed.html).
|
||||||
|
|
||||||
|
Schemathesis generates property-based tests from your OpenAPI or GraphQL schema, using Hypothesis under the hood. Its docs [recommend the CLI for most users](https://schemathesis.readthedocs.io/en/stable/faq/), since the pytest integration has fewer features.
|
||||||
|
|
||||||
|
unittest.mock ships with Python. [Patch where an object is looked up](https://docs.python.org/3/library/unittest.mock.html), not where it's defined, and add `autospec=True` so your tests fail when the real API changes. For code you own, pytest's docs suggest you [pass dependencies in](https://docs.pytest.org/en/stable/how-to/monkeypatch.html) rather than patch them.
|
||||||
|
|
||||||
|
For HTTP, pick the mock that matches your client: responses for requests, and RESPX for HTTPX. Both raise an error on requests you didn't mock. VCR.py records real responses to a cassette file and replays them, with many clients including both. [Filter out credentials](https://vcrpy.readthedocs.io/en/latest/advanced.html) before you commit cassettes. For time, FreezeGun [freezes `datetime` and `time`](https://github.com/spulec/freezegun) at the moment you choose.
|
||||||
|
|
||||||
|
factory_boy [replaces static fixtures with factories](https://factoryboy.readthedocs.io/en/stable/) that set only the fields a test cares about, and works with Django, SQLAlchemy, and MongoDB models. Polyfactory [builds objects from type hints](https://polyfactory.litestar.dev/latest/): dataclasses, TypedDicts, Pydantic models, and more.
|
||||||
|
|
||||||
|
For the data itself, Faker [generates localized fake data](https://faker.readthedocs.io/en/master/) and comes with a pytest fixture; factory_boy uses it too. Mimesis is [fully typed and generates data from schemas](https://mimesis.name/latest/about.html), in many languages.
|
||||||
|
|
||||||
|
Coverage.py measures which lines your tests run. Run pytest under it with [`coverage run -m pytest`](https://coverage.readthedocs.io/en/latest/), which its docs say is enough for most purposes, and include your tests in the measurement. It measures lines by default; add [`--branch`](https://coverage.readthedocs.io/en/latest/branch.html) to see which branches never ran.
|
||||||
|
|
||||||
|
Write each test so it [runs in any order](https://www.selenium.dev/documentation/test_practices/discouraged/test_dependency/), without relying on other tests. A [flaky test](https://docs.pytest.org/en/stable/explanation/flaky.html) usually means state the test doesn't control, and random data is one such state: seed Faker, Mimesis, factory_boy, and Polyfactory so a [failing build reproduces](https://factoryboy.readthedocs.io/en/stable/).
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
Need a Python text processing library? charset-normalizer reads text in unknown encodings, RapidFuzz does fuzzy matching, and pyparsing builds parsers.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Bytes in an unknown encoding: charset-normalizer
|
||||||
|
- Fuzzy matching against many strings: RapidFuzz
|
||||||
|
- Parsers for your own grammar: pyparsing, or parsy for small languages
|
||||||
|
- Text garbled by a wrong decode (mojibake): ftfy
|
||||||
|
- Diffs, or close matches without a dependency: difflib
|
||||||
|
- ASCII-art banners: pyfiglet
|
||||||
|
- Translations, plus localized dates and numbers: Babel
|
||||||
|
- Syntax highlighting: Pygments
|
||||||
|
- Splitting and formatting SQL: sqlparse
|
||||||
|
- Phone numbers: phonenumbers
|
||||||
|
- URL slugs: python-slugify, or Unidecode for ASCII transliteration alone
|
||||||
|
- Short public IDs: shortuuid, or Sqids to encode integer IDs
|
||||||
|
|
||||||
|
Decode with the encoding you know before you guess one. ftfy's docs say to [assume UTF-8](https://ftfy.readthedocs.io/en/latest/avoid.html#assume-utf-8) until proven otherwise, and charset-normalizer's FAQ calls detection [the last resort](https://charset-normalizer.readthedocs.io/en/latest/community/faq.html#should-i-bother-using-detection). For bytes you have no clue about, charset-normalizer is the detector [requests installs](https://github.com/psf/requests/blob/main/pyproject.toml). Call `from_bytes()` or `from_path()`, then `.best()` to [get the most probable result](https://charset-normalizer.readthedocs.io/en/latest/user/advanced_search.html). Its `detect()` is [backward compatible with chardet's](https://charset-normalizer.readthedocs.io/en/latest/user/getstarted.html), so switching is one import. If you stay on chardet, read large files and streams with [`UniversalDetector`](https://chardet.readthedocs.io/en/latest/usage.html).
|
||||||
|
|
||||||
|
ftfy fixes text that was already decoded wrong, and it [doesn't take bytes](https://ftfy.readthedocs.io/en/latest/detect.html), so it's no replacement for a detector. Call `ftfy.fix_text()`, [the function you'll use most](https://ftfy.readthedocs.io/en/latest/explain.html#ftfy.fix_text). Every fix is on by default, so [check which ones fit your text](https://ftfy.readthedocs.io/en/latest/config.html).
|
||||||
|
|
||||||
|
RapidFuzz is a fast string matching library with a C++ core. To compare one string with a list, use the process module's `process.extractOne()` or `process.cdist()`, which is [faster than calling the scorers yourself](https://github.com/rapidfuzz/RapidFuzz). It doesn't lowercase or strip punctuation for you: pass `processor=utils.default_process` when case and punctuation shouldn't count.
|
||||||
|
|
||||||
|
difflib ships with Python, so it works where you can't add a dependency. `get_close_matches()` returns the [good enough matches](https://docs.python.org/3/library/difflib.html#difflib.get_close_matches) for a word, and its diffs come out in unified, context, or HTML format.
|
||||||
|
|
||||||
|
pyparsing lets you [build the grammar in Python code](https://github.com/pyparsing/pyparsing) instead of regular expressions. Since a grammar can accept invalid input, its docs say to use it on [input you assume is well-formatted](https://pyparsing-docs.readthedocs.io/en/latest/HowToUsePyparsing.html#usage-notes). Its [best practices](https://github.com/pyparsing/pyparsing/blob/master/pyparsing/ai/best_practices.md) say to write the grammar out in BNF first. When the grammar is recursive, call `enable_packrat()` right after the import, and give results names to the fields you read back.
|
||||||
|
|
||||||
|
parsy combines small parsers into bigger ones. Its docs say it [excels at small languages](https://parsy.readthedocs.io/en/latest/overview.html) and is easy to read, but to look elsewhere when you need speed or good error messages. For more complex parsers, the [`@generate` decorator](https://parsy.readthedocs.io/en/latest/ref/generating.html) is both more readable and more powerful.
|
||||||
|
|
||||||
|
Babel does two jobs: gettext message catalogs, and [CLDR locale data](https://babel.pocoo.org/en/latest/intro.html) for localized dates, numbers, and names. Run the catalog steps with the `pybabel` command: [extract, init, update, and compile](https://babel.pocoo.org/en/latest/cmdline.html). Keep times in UTC, and [convert to the user's time zone](https://babel.pocoo.org/en/latest/dates.html#time-zone-support) only for input and display.
|
||||||
|
|
||||||
|
Pygments turns code into HTML, LaTeX, ANSI, and more, as a command-line tool or a library. For HTML, it writes CSS classes instead of inline styles, so [generate the stylesheet](https://pygments.org/docs/quickstart/#example) with `HtmlFormatter().get_style_defs()`. Pick the lexer by name or file name, and [guess it](https://pygments.org/docs/quickstart/#guessing-lexers) only when you don't know the language. Pygments [doesn't guarantee how long it runs](https://pygments.org/docs/security/), so on user input, run it with a short timeout and cap how many run at once.
|
||||||
|
|
||||||
|
sqlparse is a [non-validating SQL parser](https://sqlparse.readthedocs.io/en/latest/): it splits scripts into statements, formats them, and walks their tokens, without assuming a SQL dialect. Use `split()`, `format()`, and `parse()`. On SQL from untrusted sources, [keep its grouping limits](https://sqlparse.readthedocs.io/en/latest/api.html#security-and-performance-considerations) as they are.
|
||||||
|
|
||||||
|
phonenumbers is a Python port of Google's libphonenumber. Pass `parse()` the region the number was dialed from, unless it's in E.164 format. Then [check it's possible and valid](https://github.com/daviddrysdale/python-phonenumbers) with `is_possible_number()` and `is_valid_number()`.
|
||||||
|
|
||||||
|
python-slugify makes URL slugs, and Unidecode turns Unicode text into ASCII. Their licenses differ. [Unidecode is GPL](https://github.com/avian2/unidecode). python-slugify's own code is MIT, and by default it runs on text-unidecode, which offers the Artistic license or GPL. But python-slugify [switches to Unidecode](https://github.com/un33k/python-slugify) whenever it's installed.
|
||||||
|
|
||||||
|
shortuuid turns UUIDs into [short IDs for users to see](https://github.com/skorokithakis/shortuuid), and leaves out look-alike characters like l, 1, I, O, and 0. Sqids turns database keys and other integers into short IDs. Anyone can [decode them back into numbers](https://sqids.org/faq#not-recommended), so keep them away from sensitive data and user IDs. To check an ID is the canonical one, [re-encode the decoded numbers](https://sqids.org/faq#valid-ids) and compare.
|
||||||
|
|
||||||
|
Store what these libraries generate, or pin their versions, since slugs and IDs can change between releases. python-slugify says to [pin the package and its backend](https://github.com/un33k/python-slugify), and Unidecode says to [store each slug once or lock the version](https://github.com/avian2/unidecode). Sqids says to [pass your own blocklist](https://sqids.org/faq#future-blocklist), even one identical to the default.
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
For a Django project, pick Django REST framework or Django Ninja. Outside Django, build on FastAPI, a Python API framework based on type hints.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Django, with CRUD endpoints over your models: Django REST framework
|
||||||
|
- Django, with endpoints declared by type hints: Django Ninja
|
||||||
|
- A new service outside Django: FastAPI
|
||||||
|
- GraphQL over Django models: Strawberry GraphQL Django
|
||||||
|
- Django, with msgspec, attrs, or dataclass schemas: django-modern-rest
|
||||||
|
- A Flask app: APIFlask
|
||||||
|
- An OpenAPI spec written before the code: Connexion
|
||||||
|
- GraphQL on FastAPI, Flask, or another framework: Strawberry
|
||||||
|
- RPC between services in any language: gRPC
|
||||||
|
|
||||||
|
Django REST framework adds serializers and a [browsable API](https://www.django-rest-framework.org/) to Django. For CRUD over your models, [group the views in a `ModelViewSet` and register it with a router](https://www.django-rest-framework.org/tutorial/quickstart/). For an OpenAPI schema, its docs [recommend drf-spectacular](https://www.django-rest-framework.org/topics/documenting-your-api/#drf-spectacular). By default, the API [allows unrestricted access](https://www.django-rest-framework.org/api-guide/permissions/#setting-the-permission-policy), so set a policy in `DEFAULT_PERMISSION_CLASSES`.
|
||||||
|
|
||||||
|
Django Ninja is [heavily inspired by FastAPI](https://django-ninja.dev/motivation/) and works with Django's ORM, URLs, views, and auth. You write each endpoint as a function with type hints, and it [generates the OpenAPI docs](https://django-ninja.dev/) from them. Run async views on [an ASGI server](https://django-ninja.dev/guides/async-support/). If you use Django's cookie-based auth, [keep Django's CSRF protection on](https://django-ninja.dev/reference/csrf/#use-djangos-built-in-csrf-protection).
|
||||||
|
|
||||||
|
FastAPI builds [request validation and OpenAPI docs from standard type hints](https://fastapi.tiangolo.com/), on top of Pydantic and Starlette. Install `fastapi[standard]`, which brings Uvicorn and the `fastapi` command: run `fastapi dev` while you work and [`fastapi run` in production](https://fastapi.tiangolo.com/fastapi-cli/).
|
||||||
|
|
||||||
|
Strawberry GraphQL Django [builds types from your Django models](https://strawberry.rocks/docs/django), while Strawberry alone [only provides a GraphQL view](https://strawberry.rocks/docs/integrations/django) for Django. Add its [query optimizer](https://strawberry.rocks/docs/django/guide/optimizer), which calls `select_related()` and `prefetch_related()` for you to avoid N+1 queries.
|
||||||
|
|
||||||
|
django-modern-rest [drops into an existing Django app](https://django-modern-rest.readthedocs.io/en/latest/) and validates both requests and responses against your schemas: msgspec, Pydantic, attrs, dataclasses, and more. Its docs [recommend always installing msgspec](https://django-modern-rest.readthedocs.io/en/latest/pages/getting-started.html#installation) to parse JSON, even when your schemas are Pydantic models.
|
||||||
|
|
||||||
|
APIFlask is a [thin wrapper on Flask](https://apiflask.com/migrations/flask/): swap `Flask` for `APIFlask` and `Blueprint` for `APIBlueprint`, and your [Flask extensions keep working](https://apiflask.com/comparison/#apiflask-vs-fastapi). Declare each endpoint's input and output with `@app.input()` and `@app.output()`, as [marshmallow schemas or Pydantic models](https://apiflask.com/), and APIFlask generates the OpenAPI docs from them.
|
||||||
|
|
||||||
|
Connexion is spec-first: you write the OpenAPI spec, and Connexion [routes and validates requests against it](https://connexion.readthedocs.io/en/latest/#why-connexion), so server and client can be built in parallel. FastAPI goes the other way, and its maintainer says it's [not meant for writing the schema first](https://github.com/fastapi/fastapi/discussions/6169). Start a new project on [`AsyncApp`](https://connexion.readthedocs.io/en/latest/quickstart.html#creating-your-application), or on `FlaskApp` to keep the Flask ecosystem.
|
||||||
|
|
||||||
|
Strawberry builds a GraphQL schema from [dataclasses and type hints](https://strawberry.rocks/docs), and FastAPI's docs [recommend it for GraphQL](https://fastapi.tiangolo.com/how-to/graphql/#graphql-with-strawberry). Before you deploy, [turn off GraphiQL and introspection](https://strawberry.rocks/docs/operations/deployment), and add the [security extensions](https://strawberry.rocks/docs/operations/deployment#security-extensions) that limit query depth, aliases, and tokens.
|
||||||
|
|
||||||
|
gRPC has you [define a service once in a `.proto` file](https://grpc.io/docs/languages/python/basics/) and generate its clients and servers in any language gRPC supports. Install `grpcio` and `grpcio-tools`, and [compile the `.proto` file into Python code with `grpc_tools.protoc`](https://grpc.io/docs/languages/python/quickstart/). [Use TLS](https://grpc.io/docs/guides/auth/) to authenticate the server and encrypt the traffic.
|
||||||
|
|
||||||
|
Keep separate models for what an endpoint takes in and what it sends back. FastAPI [filters the response through the output model](https://fastapi.tiangolo.com/tutorial/response-model/#add-an-output-model), so fields the output model leaves out, like a password, never reach the client. APIFlask [recommends separate input and output schemas](https://apiflask.com/schema/#marshmallow) too.
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
Django serves static files only in development, but django-storages puts them and user uploads on S3 or other clouds. Django Compressor bundles CSS and JS.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Static files and user uploads on S3, Google Cloud Storage, or Azure: django-storages
|
||||||
|
- CSS and JavaScript combined and minified from your templates: Django Compressor
|
||||||
|
|
||||||
|
django-storages is [a collection of storage backends for Django](https://django-storages.readthedocs.io/en/latest/): Amazon S3, Google Cloud Storage, Azure Storage, and more. Install the extra for your backend, like [`django-storages[s3]`](https://django-storages.readthedocs.io/en/latest/backends/amazon-S3.html#installation). Configure it through Django's [`STORAGES` setting](https://docs.djangoproject.com/en/stable/ref/settings/#storages), with the backend's settings under `OPTIONS`. The `default` key stores user uploads, and a `staticfiles` key has `collectstatic` [put your static files on S3](https://django-storages.readthedocs.io/en/latest/backends/amazon-S3.html#configuration-settings) too. Django's security docs say to [serve user uploads from a separate domain](https://docs.djangoproject.com/en/stable/topics/security/#user-uploaded-content), not a subdomain of your site.
|
||||||
|
|
||||||
|
Django Compressor [processes, combines, and minifies](https://github.com/django-compressor/django-compressor) the CSS and JavaScript in your Django templates into cacheable static files. It also supports compilers like Sass and LESS. Its maintainer sees it as [an alternative for people who don't want to keep up](https://github.com/django-compressor/django-compressor/discussions/1095) with JavaScript build tools and want one easy setup for development and production. Add `compressor` to `INSTALLED_APPS` and [its finder](https://django-compressor.readthedocs.io/en/stable/quickstart.html) to `STATICFILES_FINDERS`. Then wrap your `<link>` and `<script>` tags in `{% compress css %}` and `{% compress js %}` blocks. In production, its docs [strongly recommend a real cache backend](https://django-compressor.readthedocs.io/en/stable/usage.html). On several servers, or with compressed files on a CDN, use [offline compression](https://django-compressor.readthedocs.io/en/stable/scenarios.html#offline-compression): set `COMPRESS_OFFLINE` and run `manage.py compress` when you deploy. Jinja2 templates work through [its extension](https://django-compressor.readthedocs.io/en/stable/jinja2.html).
|
||||||
|
|
||||||
|
django-storages and Django Compressor work together. Django Compressor's docs [point to django-storages](https://django-compressor.readthedocs.io/en/stable/remote-storages.html#django-storages) for pushing compressed files to S3, and it saves them to the storage you define under [the `compressor` alias](https://django-compressor.readthedocs.io/en/stable/settings.html#django.conf.settings.COMPRESS_STORAGE_ALIAS) in `STORAGES`.
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
When you want an ORM, an admin, and auth built in, use Django. Flask, a smaller Python web framework, gives you a core you extend.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- ORM, admin, and auth built in: Django
|
||||||
|
- A small core you extend with your own picks: Flask
|
||||||
|
- One file with no dependencies: Bottle
|
||||||
|
- An app that starts small and may grow large: Pyramid
|
||||||
|
- HTML from the server with HTMX: FastHTML
|
||||||
|
- A minimal async toolkit to build on: Starlette
|
||||||
|
- Long polling, WebSockets, and other long-lived connections: Tornado
|
||||||
|
- Async with sessions, caching, and ORM integration built in: Litestar
|
||||||
|
- Frontend and backend in pure Python: Reflex
|
||||||
|
|
||||||
|
Django comes with [a full stack for convenience](https://docs.djangoproject.com/en/stable/misc/design-philosophies/), and its pieces stay independent where possible. You describe your database layout as models in Python, and Django [builds an admin interface from them](https://docs.djangoproject.com/en/stable/intro/overview/). [User authentication](https://docs.djangoproject.com/en/stable/topics/auth/) with accounts, groups, and permissions is built in too. A project holds your settings and [one or more apps](https://docs.djangoproject.com/en/stable/intro/tutorial01/#creating-the-polls-app), and an app can move between projects. Before you deploy, run [`manage.py check --deploy`](https://docs.djangoproject.com/en/stable/howto/deployment/checklist/#run-manage-py-check-deploy) against your production settings. Async views work under WSGI, but for slow streaming and long polling, [deploy Django under ASGI](https://docs.djangoproject.com/en/stable/topics/async/#async-views).
|
||||||
|
|
||||||
|
Flask keeps [the core simple but extensible](https://flask.palletsprojects.com/en/stable/design/#what-does-micro-mean): it doesn't pick your database or form library, and extensions add them. Create the app in an [application factory](https://flask.palletsprojects.com/en/stable/patterns/appfactories/) and bind each extension with `init_app`, so tests can build instances with their own settings. Flask runs each async view on a separate thread, not an event loop, which [costs performance compared with ASGI frameworks](https://flask.palletsprojects.com/en/stable/design/#async-await-and-asgi-support). For a mainly async codebase, its docs point you to an async-first framework.
|
||||||
|
|
||||||
|
Bottle is [a single file module](https://bottlepy.org/docs/stable/) with no dependencies outside the standard library. Its FAQ pitches it for [prototyping, weekend projects, and small applications](https://bottlepy.org/docs/stable/faq.html#is-bottle-suitable-for-complex-applications), and suggests a full-stack framework like Django when you have tight deadlines. In production, have [Gunicorn or another WSGI server load your app](https://bottlepy.org/docs/stable/deployment.html) instead of calling `run()`.
|
||||||
|
|
||||||
|
Pyramid is built so you [don't have to rewrite a small app in another framework](https://docs.pylonsproject.org/projects/pyramid/en/latest/narr/introduction.html#what-makes-pyramid-unique) when it gets too big. Its core maps URLs to code, handles security, and serves static assets, and it makes no assertions about which database or template system you use. Start from [its cookiecutter](https://docs.pylonsproject.org/projects/pyramid/en/latest/narr/project.html), which asks for your template language, persistence, and URL mapping. Deploy with the `production.ini` it generates, which turns off the interactive debugger.
|
||||||
|
|
||||||
|
FastHTML is [designed to create hypermedia applications](https://www.fastht.ml/about/tech): it returns HTML from the server, the approach HTMX uses, and you often won't write any JavaScript at all. It's built on Starlette and Uvicorn. Its docs say [not to assume other frameworks' best practices apply](https://www.fastht.ml/docs/ref/best_practice.html): let the function name define each route, and use only GET and POST.
|
||||||
|
|
||||||
|
Starlette is [a lightweight ASGI framework/toolkit](https://starlette.dev/#framework-or-toolkit): use it as a complete framework, or take any of its components on their own. Install an ASGI server such as Uvicorn next to it, and pick [any async database library](https://starlette.dev/database/) you like. Keep configuration [in environment variables or a `.env` file](https://starlette.dev/config/) that you don't commit.
|
||||||
|
|
||||||
|
Tornado [isn't based on WSGI](https://www.tornadoweb.org/en/stable/#threads-and-wsgi) and typically runs one thread per process, with its own web framework and HTTP server used together. Its non-blocking I/O makes it a fit for [long polling, WebSockets, and other long-lived connections](https://www.tornadoweb.org/en/stable/). Hand blocking code to `run_in_executor`, and [run one process per CPU](https://www.tornadoweb.org/en/stable/guide/running.html#processes-and-ports).
|
||||||
|
|
||||||
|
Litestar is [not a microframework](https://docs.litestar.dev/latest/#philosophy): it comes with ORM integration, client- and server-side sessions, and caching, though it will never have its own ORM. Class-based controllers sit at its core. Set dependencies, guards, and middleware on [any layer](https://docs.litestar.dev/latest/onboarding/flask.html), from the app down to one handler, and the setting closest to the handler wins.
|
||||||
|
|
||||||
|
Reflex builds the [frontend, backend, and database in pure Python](https://reflex.dev/docs/getting-started/introduction/). It [compiles your UI to a React frontend](https://reflex.dev/docs/advanced-onboarding/how-reflex-works/), runs your state handlers on the server, and syncs the two over WebSockets. Change state only through event handlers on your `State` class. In production, the Reflex team runs Redis as the state manager. Keep auth data and other sensitive state in [backend-only vars](https://reflex.dev/docs/vars/base-vars/#backend-only-vars).
|
||||||
|
|
||||||
|
Don't deploy on the development server: [Flask](https://flask.palletsprojects.com/en/stable/deploying/) and [Django](https://docs.djangoproject.com/en/stable/howto/deployment/checklist/#switch-away-from-manage-py-runserver) both tell you to switch to a production server, and to turn debug mode off.
|
||||||
|
|
||||||
|
For a site with forms and logins, turn on CSRF protection. Django's [CSRF middleware is on by default](https://docs.djangoproject.com/en/stable/howto/csrf/). Tornado, Pyramid, and Litestar have it as a setting you turn on. Flask [leaves it to a form library](https://flask.palletsprojects.com/en/stable/web-security/#cross-site-request-forgery-csrf), and Starlette to third-party middleware.
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
To crawl a whole site, use Scrapy as your Python web scraping library; to feed pages to an LLM, Crawl4AI; to pull an article's main text, Trafilatura.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- Crawling a whole site, or many sites: Scrapy
|
||||||
|
- Web pages as clean Markdown for an LLM: Crawl4AI
|
||||||
|
- An article's main text and metadata: Trafilatura
|
||||||
|
- An AI agent that does a task in the browser for you: Browser Use
|
||||||
|
- Browser automation in your own code, with plain-language steps: Stagehand
|
||||||
|
- A browser agent that picks each action from a table of page elements: jev-ultrafast
|
||||||
|
- RSS, Atom, and JSON feeds: feedparser
|
||||||
|
- A whole HTML page as Markdown text: html2text
|
||||||
|
|
||||||
|
Scrapy is a framework [for crawling websites and extracting structured data](https://docs.scrapy.org/en/latest/intro/overview.html) from them, and it sends requests asynchronously. Create a project with `scrapy startproject` and [write your spiders in it](https://docs.scrapy.org/en/latest/intro/tutorial.html). When a page loads its data with JavaScript, [find the request that returns the data](https://docs.scrapy.org/en/latest/topics/dynamic-content.html) and send it yourself. Use a headless browser only when that fails. To run your spiders in production, deploy them to [Scrapyd or Zyte Scrapy Cloud](https://docs.scrapy.org/en/latest/topics/deploy.html).
|
||||||
|
|
||||||
|
Crawl4AI [turns websites into clean, LLM-ready Markdown](https://docs.crawl4ai.com/). It runs Chromium headless by default, so run `crawl4ai-setup` once to install the browser. Open an `AsyncWebCrawler` in an `async with` block and [call `arun()` for each URL](https://docs.crawl4ai.com/core/quickstart/). When pages share one layout, extract the data with CSS or XPath schemas: they [do exactly what you specify](https://docs.crawl4ai.com/extraction/no-llm-strategies/), while an LLM's output can vary. Keep [LLM extraction](https://docs.crawl4ai.com/extraction/llm-strategies/), which is slower and costlier, for content an AI has to interpret.
|
||||||
|
|
||||||
|
Trafilatura [extracts a page's main text and metadata](https://trafilatura.readthedocs.io/en/latest/) and skips boilerplate like headers and footers. Call `extract()`, and pass `favor_precision=True` or `favor_recall=True` to [tune what it keeps](https://trafilatura.readthedocs.io/en/latest/usage-python.html). It works on raw HTML, so [render JavaScript pages first](https://trafilatura.readthedocs.io/en/latest/troubleshooting.html) with a browser and pass the result to `extract()`. It pairs with Scrapy: [Scrapy crawls, Trafilatura extracts](https://trafilatura.readthedocs.io/en/latest/faq.html#how-does-trafilatura-compare-to-beautifulsoup-or-scrapy).
|
||||||
|
|
||||||
|
Browser Use is an AI browser agent: you [give it a task and an LLM](https://docs.browser-use.com/open-source/quickstart), and it runs the task in a browser. Its docs say to [be specific about the actions](https://docs.browser-use.com/open-source/customize/agent/prompting-guide) you want. Pass a Pydantic model as `output_model_schema` to get structured results. Don't take the agent's word for it: `is_successful()` is [only its own assessment](https://docs.browser-use.com/open-source/customize/agent/output-format), so check important results yourself, like whether a form really got submitted. Pass logins as `sensitive_data`, so the model [sees only placeholders](https://docs.browser-use.com/open-source/examples/templates/sensitive-data) in the text it reads. Set `use_vision=False` too, or the real values can leak through screenshots.
|
||||||
|
|
||||||
|
Stagehand leaves the steps to your code: you [mix plain-language actions with regular page calls](https://docs.stagehand.dev/first-steps/introduction) in one script, and decide how much AI each step uses. Give each `act()` call [one focused action](https://docs.stagehand.dev/best-practices/prompting-best-practices), and pass credentials as variables, so their values never reach the model. To get data out, pass `extract()` a Pydantic model, and Stagehand [validates the result against it](https://docs.stagehand.dev/basics/extract).
|
||||||
|
|
||||||
|
feedparser parses RSS, Atom, and JSON feeds with [one function, `parse()`](https://feedparser.readthedocs.io/en/latest/introduction/), which takes a URL, a file, or a string. When you poll a feed, send back the [ETag and Last-Modified values](https://feedparser.readthedocs.io/en/latest/http-etag/) from the last response. Otherwise you download unchanged feeds again, and the publisher may ban you. Content feedparser marks as `text/plain` [hasn't been sanitized](https://feedparser.readthedocs.io/en/latest/html-sanitization/), so escape it before you render it.
|
||||||
|
|
||||||
|
html2text [converts a page of HTML into Markdown](https://github.com/Alir3z4/html2text), with options such as [`ignore_links`](https://github.com/Alir3z4/html2text/blob/master/docs/usage.md). It's GPL, while Trafilatura is [Apache](https://trafilatura.readthedocs.io/en/latest/).
|
||||||
|
|
||||||
|
Whatever you pick, tell sites who you are and go easy on them. Scrapy's docs say to [set `USER_AGENT` to a value that identifies you](https://docs.scrapy.org/en/latest/topics/practices.html#avoiding-getting-banned) and space out your requests. The feedparser docs say to [set the User-Agent to your app's name and URL](https://feedparser.readthedocs.io/en/latest/http-useragent/), and Trafilatura's to [throttle per domain and follow robots.txt](https://trafilatura.readthedocs.io/en/latest/downloads.html).
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
Every response your app sends should carry security headers, and a Python web security library like secure defines them once for Django, Flask, or FastAPI.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- FastAPI or Starlette: secure's ASGI middleware
|
||||||
|
- Django, Flask, and other frameworks: secure in a response hook or middleware
|
||||||
|
- Most apps: the default BALANCED preset
|
||||||
|
- Tighter CSP and stronger isolation: the STRICT preset
|
||||||
|
|
||||||
|
secure lets you [keep one `Secure` policy object instead of scattered header strings](https://typeerror.com/secure/), and applies it through ASGI middleware, WSGI middleware, or your framework's response hooks. Start with `Secure.with_default_headers()`, which matches `Preset.BALANCED`, [the recommended default for most applications](https://github.com/TypeError/secure/blob/main/docs/usage.md#start-with-a-preset). [Review the headers it generates, and move stricter only when needed](https://typeerror.com/secure/#presets): `Preset.STRICT` tightens the CSP and isolation.
|
||||||
|
|
||||||
|
The framework guide says to [prefer middleware when your framework makes it easy](https://github.com/TypeError/secure/blob/main/docs/frameworks.md#how-to-choose-an-integration-style) and you want app-wide coverage. On FastAPI, [add `SecureASGIMiddleware` with `app.add_middleware()`](https://github.com/TypeError/secure/blob/main/docs/frameworks.md#fastapi). On Flask, [call `set_headers(response)` in an `after_request` hook](https://github.com/TypeError/secure/blob/main/docs/frameworks.md#flask). On Django, write [a small Django middleware class](https://github.com/TypeError/secure/blob/main/docs/frameworks.md#django) that calls `set_headers()` and register it in `MIDDLEWARE`. For a framework the guide doesn't cover, [configure one `Secure` instance and apply it to the response as late as possible](https://github.com/TypeError/secure/blob/main/docs/frameworks.md#custom-frameworks) before it's sent.
|
||||||
|
|
||||||
|
The defaults are [a starting point, not a substitute for a review against your own app](https://typeerror.com/secure/). Adjust the Content Security Policy in particular for the scripts, styles, assets, and third-party services your app actually uses. Build it with the `ContentSecurityPolicy` builder instead of a header string, and [test a stricter policy against the real app before rollout](https://github.com/TypeError/secure/blob/main/docs/usage.md#build-an-explicit-configuration).
|
||||||
|
|
||||||
|
Security headers help the browser enforce transport, embedding, and content-loading rules, but they [don't replace output encoding, CSRF protection, authentication, or input validation](https://github.com/TypeError/secure/blob/main/docs/security_considerations.md). For web security materials beyond Python libraries, see [awesome-web-security](https://github.com/qazbnm456/awesome-web-security).
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
In production, Gunicorn serves WSGI apps like Django, and Uvicorn serves ASGI apps like FastAPI. Waitress, a pure-Python web server, runs WSGI apps on Windows.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A WSGI app (Flask, Django) on Linux or macOS: Gunicorn
|
||||||
|
- An ASGI app (FastAPI, Starlette): Uvicorn
|
||||||
|
- A WSGI app on Windows, or a server in pure Python: Waitress
|
||||||
|
- One server for both ASGI and WSGI apps, or throughput above all: Granian
|
||||||
|
- An app on Trio, or HTTP/3 from the Python server: Hypercorn
|
||||||
|
|
||||||
|
Gunicorn is a [pre-fork server](https://gunicorn.org/design/#server-model): one process manages a pool of worker processes. It's [built for Unix](https://gunicorn.org/). With its default sync workers, your proxy must [buffer slow clients](https://gunicorn.org/deploy/#nginx-configuration), or Gunicorn is open to denial-of-service attacks. Start with (2 × CPU cores) + 1 workers and [adjust under load](https://gunicorn.org/design/#how-many-workers). For long blocking calls, streaming, or WebSockets, switch to [async workers](https://gunicorn.org/design/#when-to-use-async-workers) like gevent.
|
||||||
|
|
||||||
|
Uvicorn is an [ASGI web server](https://uvicorn.dev/). In production, run it under a [process manager](https://uvicorn.dev/deployment/#using-a-process-manager). Its [built-in one](https://uvicorn.dev/deployment/#built-in) starts workers with `--workers` and restarts any that die, and its docs also cover [running Uvicorn workers under Gunicorn](https://uvicorn.dev/deployment/#gunicorn). In a container, run a single Uvicorn process and [let your orchestrator scale](https://uvicorn.dev/deployment/docker/) the number of containers.
|
||||||
|
|
||||||
|
Granian is a [Rust HTTP server](https://github.com/emmett-framework/granian) that serves ASGI, WSGI, and RSGI apps from one package. Its README [suggests it](https://github.com/emmett-framework/granian#rationale) when you care about throughput above all, and not when you want pure Python or your app relies on Trio or gevent. Pass [`--interface asgi` or `--interface wsgi`](https://github.com/emmett-framework/granian#options), since the default is RSGI. Start with [one worker per CPU core](https://github.com/emmett-framework/granian#workers-and-threads), or one per container on Docker or Kubernetes, rather than numbers suggested for other servers.
|
||||||
|
|
||||||
|
Hypercorn is an ASGI server [inspired by Gunicorn](https://hypercorn.readthedocs.io/en/latest/). It speaks HTTP/1 and HTTP/2, with WebSockets over both, and runs on [Trio](https://hypercorn.readthedocs.io/en/latest/discussion/workers.html#trio) as well as asyncio. For HTTP/3, install its [`h3` extra](https://github.com/pgjones/hypercorn). Its docs recommend setting [`server_names`](https://hypercorn.readthedocs.io/en/latest/how_to_guides/server_names.html#dns-rebinding-attacks) to the hosts you serve, to block DNS rebinding attacks.
|
||||||
|
|
||||||
|
Waitress is a [pure-Python WSGI server](https://docs.pylonsproject.org/projects/waitress/en/latest/index.html) with no dependencies outside the standard library, and it runs on both Unix and Windows. Waitress [doesn't support TLS](https://docs.pylonsproject.org/projects/waitress/en/latest/reverse-proxy.html) itself, so put a reverse proxy in front of it for HTTPS.
|
||||||
|
|
||||||
|
Put a proxy like Nginx in front: Gunicorn's docs [strongly recommend it](https://gunicorn.org/deploy/), and Uvicorn's recommend it [for resilience](https://uvicorn.dev/deployment/#running-behind-nginx). Behind a proxy, tell your server which proxies to trust for `X-Forwarded-*` headers, since any client can set them. [Uvicorn](https://uvicorn.dev/deployment/#proxies-and-forwarded-headers) and Gunicorn take `--forwarded-allow-ips`, Granian [`trusted_hosts`](https://github.com/emmett-framework/granian#proxies-and-forwarded-headers), Hypercorn [`ProxyFixMiddleware`](https://hypercorn.readthedocs.io/en/latest/how_to_guides/proxy_fix.html), and Waitress [`trusted_proxy`](https://docs.pylonsproject.org/projects/waitress/en/latest/reverse-proxy.html#passing-the-proxy-headers-to-setup-the-wsgi-environment). In Uvicorn, Gunicorn, Granian, and Waitress, trust every address with `*` only when [no client can reach the server directly](https://gunicorn.org/deploy/#nginx-configuration).
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
With Django, Channels is the pick; with Flask and Socket.IO clients, Flask-SocketIO. Standalone apps run on websockets, a Python WebSocket library.
|
||||||
|
|
||||||
|
How to choose:
|
||||||
|
|
||||||
|
- A Django project: Channels
|
||||||
|
- A Flask app with Socket.IO clients: Flask-SocketIO
|
||||||
|
- Plain WebSocket servers and clients: websockets
|
||||||
|
- A few real-time features next to a Django project: websockets as a separate server
|
||||||
|
- RPC and pub/sub over WAMP, or a Twisted app: Autobahn|Python
|
||||||
|
|
||||||
|
Channels [extends Django beyond HTTP](https://channels.readthedocs.io/en/latest/) to handle WebSockets, and it [integrates with Django's auth and sessions](https://channels.readthedocs.io/en/latest/introduction.html). Write [sync consumers by default](https://channels.readthedocs.io/en/latest/topics/consumers.html#basic-layout). Switch to async ones only when async handling helps and every library you call is async-native. Serve everything with Daphne, or [keep HTTP on your WSGI server](https://channels.readthedocs.io/en/latest/deploying.html#http-and-websocket) and send only WebSockets to Daphne.
|
||||||
|
|
||||||
|
Flask-SocketIO speaks Socket.IO, which is [not a WebSocket implementation](https://socket.io/docs/v4/#what-socketio-is-not): a plain WebSocket client can't connect to it. Pick it when your clients use [a Socket.IO client library](https://flask-socketio.readthedocs.io/en/latest/intro.html#requirements). In return, you get events, [rooms, and broadcasting](https://flask-socketio.readthedocs.io/en/latest/getting_started.html#rooms). Start the server with `socketio.run()`, since `flask run` [lacks WebSocket support](https://flask-socketio.readthedocs.io/en/latest/getting_started.html#initialization).
|
||||||
|
|
||||||
|
websockets runs on asyncio by default, which is [ideal for servers with many connections](https://websockets.readthedocs.io/en/stable/). For clients, its threading implementation is a good alternative. It [isn't an HTTP server](https://websockets.readthedocs.io/en/stable/faq/server.html#how-do-i-run-http-and-websocket-servers-on-the-same-port), so [run it as its own program](https://websockets.readthedocs.io/en/stable/deploy/index.html#how-do-i-start-a-process) that calls `serve()`, not under a WSGI or ASGI server. For a few real-time features in a Django project, like notifications, its docs find [a separate websockets server next to Django](https://websockets.readthedocs.io/en/stable/howto/django.html) well suited, where Channels means switching to a new deployment architecture.
|
||||||
|
|
||||||
|
Autobahn|Python implements [both WebSocket and WAMP, on Twisted or asyncio](https://autobahn.readthedocs.io/en/latest/). WAMP adds [RPC and pub/sub over WebSocket](https://github.com/crossbario/autobahn-python), and every WAMP client [needs a WAMP router to talk to](https://autobahn.readthedocs.io/en/latest/wamp/programming.html#wamp-programming-1). Write components with functions and decorators, [the recommended approach](https://autobahn.readthedocs.io/en/latest/wamp/programming.html#creating-components-1).
|
||||||
|
|
||||||
|
Always [secure WebSocket connections with TLS](https://websockets.readthedocs.io/en/stable/topics/security.html#encryption) in production. Any site can open a WebSocket to yours, with your users' cookies attached. If you serve private data, [restrict the allowed origins](https://channels.readthedocs.io/en/latest/topics/security.html#websockets): Channels has `AllowedHostsOriginValidator`, websockets has the [`origins` argument](https://websockets.readthedocs.io/en/stable/reference/asyncio/server.html#websockets.asyncio.server.serve), and Flask-SocketIO [allows only the same origin by default](https://flask-socketio.readthedocs.io/en/latest/deployment.html#cross-origin-controls).
|
||||||
|
|
||||||
|
Broadcasting across processes needs a message broker such as Redis. Channels' [production channel layer](https://channels.readthedocs.io/en/latest/topics/channel_layers.html#redis-channel-layer) runs on it, Flask-SocketIO [takes it as a message queue](https://flask-socketio.readthedocs.io/en/latest/deployment.html#using-multiple-workers), with sticky sessions at the load balancer, and websockets' docs [suggest it for pub/sub](https://websockets.readthedocs.io/en/stable/faq/server.html#how-do-i-send-a-message-to-a-channel-a-topic-or-some-users).
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
{
|
||||||
|
"/categories/": "/",
|
||||||
|
"/categories/ai-and-agents/pre-trained-models-and-inference/": "/categories/ai-and-agents/pre-trained-models/",
|
||||||
|
"/categories/cli-tools/productivity-tools/": "/categories/cli-tools/",
|
||||||
|
"/categories/code-analysis/code-linters/": "/categories/code-analysis/linters-and-formatters/",
|
||||||
|
"/categories/data-analysis/financial-data/": "/categories/data-ingestion-etl/financial-data/",
|
||||||
|
"/categories/distributed-computing/batch-processing/": "/categories/distributed-computing/",
|
||||||
|
"/categories/gui-development/terminal/": "/categories/cli-development/tui-frameworks/",
|
||||||
|
"/categories/web-servers/rpc/": "/categories/web-apis/rpc/"
|
||||||
|
}
|
||||||
+56
-87
File diff suppressed because it is too large
Load Diff
+142
-53
File diff suppressed because it is too large
Load Diff
+136
-91
File diff suppressed because it is too large
Load Diff
@@ -96,7 +96,6 @@
|
|||||||
</section>
|
</section>
|
||||||
{% endif %}
|
{% endif %}
|
||||||
|
|
||||||
<script type="application/json" id="filter-urls">{{ filter_urls_json | safe }}</script>
|
|
||||||
<section class="results-section" id="library-index">
|
<section class="results-section" id="library-index">
|
||||||
<div class="results-intro section-shell" data-reveal>
|
<div class="results-intro section-shell" data-reveal>
|
||||||
<div>
|
<div>
|
||||||
@@ -132,12 +131,6 @@
|
|||||||
aria-label="Search projects"
|
aria-label="Search projects"
|
||||||
/>
|
/>
|
||||||
</div>
|
</div>
|
||||||
<div class="filter-bar" aria-live="polite">
|
|
||||||
<span>Filtering for <strong class="filter-value"></strong></span>
|
|
||||||
<button class="filter-clear" aria-label="Clear filter">
|
|
||||||
Clear filter
|
|
||||||
</button>
|
|
||||||
</div>
|
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
<h2 class="sr-only">Results</h2>
|
<h2 class="sr-only">Results</h2>
|
||||||
@@ -175,7 +168,6 @@
|
|||||||
{% for entry in entries %}
|
{% for entry in entries %}
|
||||||
<tr
|
<tr
|
||||||
class="row"
|
class="row"
|
||||||
data-tags="{{ entry.categories | join('||') }}{% if entry.subcategories %}||{{ entry.subcategories | map(attribute='value') | join('||') }}{% endif %}||{{ entry.groups | join('||') }}{% if entry.source_type == 'Stdlib' %}||Stdlib{% endif %}"
|
|
||||||
tabindex="0"
|
tabindex="0"
|
||||||
aria-expanded="false"
|
aria-expanded="false"
|
||||||
aria-controls="expand-{{ loop.index }}"
|
aria-controls="expand-{{ loop.index }}"
|
||||||
@@ -221,23 +213,19 @@
|
|||||||
</td>
|
</td>
|
||||||
<td class="col-cat">
|
<td class="col-cat">
|
||||||
{% for subcat in entry.subcategories %}
|
{% for subcat in entry.subcategories %}
|
||||||
<a class="tag" href="{{ subcat.url }}" data-value="{{ subcat.value }}" data-url="{{ subcat.url }}">
|
<a class="tag" href="{{ category_urls[subcat.value.split(' > ')[0]] }}#{{ subcat.slug }}">
|
||||||
{{ subcat.name }}
|
{{ subcat.name }}
|
||||||
</a>
|
</a>
|
||||||
{% endfor %} {% for cat in entry.categories %}
|
{% endfor %} {% for cat in entry.categories %}
|
||||||
<a
|
<a
|
||||||
class="tag"
|
class="tag"
|
||||||
href="{{ category_urls[cat] }}"
|
href="{{ category_urls[cat] }}"
|
||||||
data-value="{{ cat }}"
|
|
||||||
data-url="{{ category_urls[cat] }}"
|
|
||||||
>{{ cat }}</a
|
>{{ cat }}</a
|
||||||
>
|
>
|
||||||
{% endfor %}
|
{% endfor %}
|
||||||
<a
|
<a
|
||||||
class="tag tag-group"
|
class="tag tag-group"
|
||||||
href="{{ filter_urls[entry.groups[0]] }}"
|
href="{{ filter_urls[entry.groups[0]] }}"
|
||||||
data-value="{{ entry.groups[0] }}"
|
|
||||||
data-url="{{ filter_urls[entry.groups[0]] }}"
|
|
||||||
>
|
>
|
||||||
{{ entry.groups[0] }}
|
{{ entry.groups[0] }}
|
||||||
</a>
|
</a>
|
||||||
@@ -245,8 +233,6 @@
|
|||||||
<a
|
<a
|
||||||
class="tag tag-source"
|
class="tag tag-source"
|
||||||
href="/categories/built-in/"
|
href="/categories/built-in/"
|
||||||
data-value="Stdlib"
|
|
||||||
data-url="/categories/built-in/"
|
|
||||||
>
|
>
|
||||||
Stdlib
|
Stdlib
|
||||||
</a>
|
</a>
|
||||||
|
|||||||
@@ -0,0 +1,11 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en">
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<title>Redirecting…</title>
|
||||||
|
<link rel="canonical" href="{{ target_url }}">
|
||||||
|
<script>location="{{ target_url }}"</script>
|
||||||
|
<meta http-equiv="refresh" content="0; url={{ target_url }}">
|
||||||
|
<meta name="robots" content="noindex">
|
||||||
|
<h1>Redirecting…</h1>
|
||||||
|
<a href="{{ target_url }}">Click here if you are not redirected.</a>
|
||||||
|
</html>
|
||||||
+236
-80
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user