diff --git a/.claude/skills/audit-the-list/SKILL.md b/.claude/skills/audit-the-list/SKILL.md index bb008997..ccb433e2 100644 --- a/.claude/skills/audit-the-list/SKILL.md +++ b/.claude/skills/audit-the-list/SKILL.md @@ -33,7 +33,7 @@ Run the `preview-verdicts` skill: it generates the interactive review page and d ## 5. Execute -One commit per section: body lists each removal with its reason and downloads figure; restructures, tier moves, and reorders ride the same commit. Format-only outcomes (no removals) are a single style commit. `make test` before every commit, `make build` after the last one. Generic commit helpers tend to split a section audit into structural and per-subcategory commits — if that happens, squash back to one commit per section. Done when the tree is clean, tests passed before each commit, and the build count reconciles with the adjudicated changes. +One commit per section: body lists each removal with its reason and downloads figure; restructures, tier moves, and reorders ride the same commit. A restructure that renames or dissolves a category or subcategory adds `"old path": "new path"` to `website/data/redirects.json` in that commit, so the old URL keeps its search ranking; the build fails when a redirect target no longer exists, so a later rename also updates the entries pointing at the renamed path. Format-only outcomes (no removals) are a single style commit. `make test` before every commit, `make build` after the last one. Generic commit helpers tend to split a section audit into structural and per-subcategory commits — if that happens, squash back to one commit per section. Done when the tree is clean, tests passed before each commit, and the build count reconciles with the adjudicated changes. ## 6. Record diff --git a/.gitignore b/.gitignore index ba307609..de692147 100644 --- a/.gitignore +++ b/.gitignore @@ -13,6 +13,8 @@ __pycache__/ website/output/ website/data/* !website/data/pypi_name_overrides.json +!website/data/redirects.json +!website/data/category_intros/ # agents .playwright-cli/ diff --git a/AGENTS.md b/AGENTS.md index 9103b0f8..01e007a9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,7 +9,15 @@ An opinionated guide to the best Python frameworks, libraries, and tools. [CONTRIBUTING.md](CONTRIBUTING.md) holds the admission rules, quality requirements, rejection rules, entry format, and ordering. Apply it whenever adding or removing an entry — direct commits included, not only PR reviews. - Every keep/drop reason must be verified against current online data at decision time — download counts, repo activity and archived status, PyPI metadata, project docs. Judging tiers — obvious choice vs challenger — also requires WebSearch evidence (adoption trajectory, community sentiment), not download counts alone. Training-data recollections are not evidence; label anything unverifiable as a judgment call. +- Checks stay read-only: judge whether a listed project installs or works from its PyPI metadata (`requires_python`, classifiers, wheel tags) and issue tracker, never by installing or running it, since that runs untrusted code. - A download count is not automatically independent demand. When one listed entry depends on another, check `requires_dist` on PyPI before citing the depended-on entry's count: mkdocs-material hard-depends on mkdocs, so mkdocs' figure exceeded mkdocs-material's by only about one percent and nearly all of it was mkdocs-material pulling it in. - One entry per commit when adding or deleting entries. Exceptions: a prune sweep is one commit per section, its body listing each removal with its reason; format, wording, or categorization changes may be bundled. Cross-section re-homes ride the originating audit's commit (both sides of the move in one diff). - Resources sections are not project entries: out of audit scope, and the website never parses them. - Sponsor placement never influences which projects get listed — see [SPONSORSHIP.md](SPONSORSHIP.md). + +## Website + +- Model the layout on https://www.placestoread.xyz: the whole list on one page, dense rows, a row that expands inline with its content aligned to the Name column, sorting by column headers, full-text search, tags that filter, a divider under the header row instead of a strong top border, a solid dark footer, and minimal decoration. Keep that model (no card grid, modal details, or pagination) unless the maintainer asks, and keep green out of the palette. +- `.section-shell` (`--shell-max: 84rem`) is the only width cap: sections, tables, rows, CTAs, and paragraphs stay uncapped, since the maintainer reads on wide screens and prose-width advice like 65-75ch doesn't apply. Remove inner `max-width` rules you come across, and ask with a concrete reason before adding one. +- Pick font sizes one step larger than feels right (`--text-base` over `--text-sm`, 1.75rem over 1.5rem for a heading), and shrink an existing size only when the maintainer asks: they asked for bigger type 11 times across 8 sessions and never for smaller. +- When you style one link, tag, or label, give its peers the same style in the same change: hero, nav, footer, project, and sponsor links share hover styles, and tag variants build on `.tag`. diff --git a/CLAUDE.md b/CLAUDE.md index 2a599c49..22541c01 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -9,7 +9,15 @@ An opinionated guide to the best Python frameworks, libraries, and tools. [CONTRIBUTING.md](CONTRIBUTING.md) holds the admission rules, quality requirements, rejection rules, entry format, and ordering. Apply it whenever adding or removing an entry — direct commits included, not only PR reviews. - Every keep/drop reason must be verified against current online data at decision time — download counts, repo activity and archived status, PyPI metadata, project docs. Judging tiers — obvious choice vs challenger — also requires WebSearch evidence (adoption trajectory, community sentiment), not download counts alone. Training-data recollections are not evidence; label anything unverifiable as a judgment call. +- Checks stay read-only: judge whether a listed project installs or works from its PyPI metadata (`requires_python`, classifiers, wheel tags) and issue tracker, never by installing or running it, since that runs untrusted code. - A download count is not automatically independent demand. When one listed entry depends on another, check `requires_dist` on PyPI before citing the depended-on entry's count: mkdocs-material hard-depends on mkdocs, so mkdocs' figure exceeded mkdocs-material's by only about one percent and nearly all of it was mkdocs-material pulling it in. - One entry per commit when adding or deleting entries. Exceptions: a prune sweep is one commit per section, its body listing each removal with its reason; format, wording, or categorization changes may be bundled. Cross-section re-homes ride the originating audit's commit (both sides of the move in one diff). - Resources sections are not project entries: out of audit scope, and the website never parses them. - Sponsor placement never influences which projects get listed — see [SPONSORSHIP.md](SPONSORSHIP.md). + +## Website + +- Model the layout on https://www.placestoread.xyz: the whole list on one page, dense rows, a row that expands inline with its content aligned to the Name column, sorting by column headers, full-text search, tags that filter, a divider under the header row instead of a strong top border, a solid dark footer, and minimal decoration. Keep that model (no card grid, modal details, or pagination) unless the maintainer asks, and keep green out of the palette. +- `.section-shell` (`--shell-max: 84rem`) is the only width cap: sections, tables, rows, CTAs, and paragraphs stay uncapped, since the maintainer reads on wide screens and prose-width advice like 65-75ch doesn't apply. Remove inner `max-width` rules you come across, and ask with a concrete reason before adding one. +- Pick font sizes one step larger than feels right (`--text-base` over `--text-sm`, 1.75rem over 1.5rem for a heading), and shrink an existing size only when the maintainer asks: they asked for bigger type 11 times across 8 sessions and never for smaller. +- When you style one link, tag, or label, give its peers the same style in the same change: hero, nav, footer, project, and sponsor links share hover styles, and tag variants build on `.tag`. diff --git a/DESIGN.md b/DESIGN.md index 1e678f2b..56ea32bc 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -140,7 +140,7 @@ Hard-won sizing rules (do not relax): Depth comes from **tonal layers**, not heavy shadows. - The page is a quiet warm canvas (`--bg-page`). The content shell is slightly brighter paper (`--bg-paper`). The sponsor band, CTA backgrounds, and inline decorative blocks step up to `--bg-paper-strong`. -- The hero is the one place that uses real atmosphere: subtle grid, slow sheen, warm radial gradients on a dark earthy ground (`--hero-bg-start` → `--hero-bg-mid` → `--hero-bg-end`). The sheen and any other motion respect `prefers-reduced-motion`. +- The hero and the category guide band are the only places with real atmosphere: warm gradients on a dark earthy ground (`--hero-bg-start` → `--hero-bg-mid` → `--hero-bg-end`). Only the hero adds the subtle grid and slow sheen. The sheen and any other motion respect `prefers-reduced-motion`. - The footer is a single tonal block in `--footer-bg`, no internal gradients. - Two depth treatments are allowed and only these two. The search input combines a 1px inset highlight (`--search-inset`) with a soft warm drop shadow (`--search-shadow`, intensified by `--search-focus-shadow` on focus). The primary CTA button (`.hero-action-primary`) carries a warm drop shadow for press affordance. Both shadows are soft, warm-tinted, and tied to interactive elements. No new drop shadows on cards, panels, rows, or static decoration. - No glassmorphism as default decoration. @@ -160,8 +160,11 @@ The shape language is overwhelmingly **pill on small, zero radius on large**. The component vocabulary is small and table-led. Source of truth: `website/static/style.css`. - **Table-driven index** (the hero of the page). Sticky header, sortable columns, click-to-expand rows that indent under the Name column. Modeled on placestoread.xyz. Not a card grid. +- **Group rows**. Section, subcategory, and group pages list rows in README order. Section pages add one H2 row per use case; group pages add one per section, linking to that section's page. A note replaces the sort arrow until someone sorts a column, which flattens the list and hides the group rows. - **Filter tags** (`.tag`). `--accent-soft` background with `--accent-deep` text. Pill shape. Hover swaps to `--highlight` background with `--tag-hover-border` border and ink text. Active state uses the warm `--tag-active-start` → `--tag-active-end` gradient with hero-ink text. Tag variants (`tag-group`, `tag-source`) inherit the base `.tag` style today and differ only at narrow widths (`tag-group` hides under 960px). Add a new variant only when a real visual difference is needed. - **Hero**. Magazine-cover headline, dark earthy ground, kicker and proof microcopy, primary CTA button using `--hero-btn-start` / `--hero-btn-end`. Subtle grid plus slow sheen. Respects `prefers-reduced-motion`. +- **Jump links**. One per use case in a category hero, plus one for the guide. Underlined like the intro's links, with a trailing ↓. Never tag-styled: tags filter the table or open another page, jump links only scroll. +- **Guide band**. Everything in a category intro after its "How to choose" list. It sits below the table on the hero's dark gradient, with its H2 above the text. - **Sponsor band**. Sits in the README header on `--bg-paper-strong`. Editorial layout, not a logo wall. Sponsor links share the global accent treatment. - **CTA**. Warm `--cta-bg`, full-bleed within shell. The button itself uses accent tokens. - **Footer**. Dark warm charcoal, part of the same system. Footer links share the global hover and focus treatment. diff --git a/website/build.py b/website/build.py index 67346cec..27b9adc0 100644 --- a/website/build.py +++ b/website/build.py @@ -13,11 +13,14 @@ from typing import TypedDict from fetch_pypi_downloads_via_clickpy import OVERRIDES_FILE, normalize from jinja2 import Environment, FileSystemLoader -from readme_parser import AlsoSee, ParsedGroup, ParsedSection, parse_readme, parse_sponsors, slugify +from markdown_it import MarkdownIt +from markdown_it.tree import SyntaxTreeNode +from readme_parser import AlsoSee, ParsedGroup, ParsedSection, parse_readme, parse_sponsors, render_inline_text, slugify GITHUB_REPO_URL_RE = re.compile(r"^https?://github\.com/([^/]+/[^/]+?)(?:\.git)?/?$") MARKDOWN_LINK_RE = re.compile(r"\[([^\]]+)\]\(([^)\s]+)\)") BULLET_LINE_RE = re.compile(r"^\s*-\s") +ANCHOR_LINK_ATTRS_RE = re.compile(r'href="(#[^"]*)" target="_blank" rel="noopener"') SITE_URL = "https://awesome-python.com/" SITEMAP_URL = f"{SITE_URL}sitemap.xml" SITEMAP_NS = "http://www.sitemaps.org/schemas/sitemap/0.9" @@ -64,6 +67,13 @@ class TemplateEntry(TypedDict): also_see: list[AlsoSee] +class EntryGroup(TypedDict): + name: str # empty for a page with a single unnamed group + slug: str + url: str # links the group heading to its own page, empty when it has none + entries: list[TemplateEntry] + + class SyntheticCategory(TypedDict): name: str slug: str @@ -220,7 +230,7 @@ def category_meta_title(name: str, parent_name: str | None = None) -> str: if len(title) <= 60: return title return f"{name} - Awesome Python" - title = f"{name} Python Libraries - Awesome Python" + title = f"Python {name} Libraries - Awesome Python" if len(title) <= 60: return title return f"{name} - Awesome Python" @@ -235,6 +245,59 @@ def category_meta_description(name: str, entry_count: int, description: str, par return f"{count_sentence} Part of the Awesome Python catalog." +def load_category_intro(path: Path) -> tuple[str, str, str]: + """Render a category intro file to HTML, split at the end of its "How to choose:" list. + + Returns the part shown above the table, the guide shown below it, and the + first paragraph as plain text for the meta description. A file without the + list keeps everything above the table. Returns empty strings if the category + has no intro file. + """ + if not path.exists(): + return "", "", "" + md = MarkdownIt("commonmark") + tokens = md.parse(path.read_text(encoding="utf-8")) + for token in tokens: + for child in token.children or []: + if child.type == "link_open": + child.attrSet("target", "_blank") + child.attrSet("rel", "noopener") + lead = next(node for node in SyntaxTreeNode(tokens).children if node.type == "paragraph") + split_at = len(tokens) + for i, token in enumerate(tokens): + if token.type == "inline" and token.level == 1 and token.content == "How to choose:" and i + 2 < len(tokens) and tokens[i + 2].type == "bullet_list_open": + split_at = next(j for j in range(i + 3, len(tokens)) if tokens[j].type == "bullet_list_close" and tokens[j].level == 0) + 1 + break + render = md.renderer.render + return render(tokens[:split_at], md.options, {}), render(tokens[split_at:], md.options, {}), render_inline_text(lead.children[0].children) + + +def group_section_entries(section: ParsedSection, entries_by_key: dict[tuple[str, str], TemplateEntry]) -> list[EntryGroup]: + """Group a section's entries by use case (subcategory), both in README order.""" + groups: dict[str, EntryGroup] = {} + for parsed in section["entries"]: + name = parsed["subcategory"] + slug = slugify(name) if name else "" + group = groups.setdefault(name, EntryGroup(name=name, slug=slug, url=subcategory_path(section["slug"], slug) if name else "", entries=[])) + group["entries"].append(entries_by_key[(parsed["url"], parsed["name"])]) + return list(groups.values()) + + +def group_entries_by_section(sections: Sequence[ParsedSection], entries_by_key: dict[tuple[str, str], TemplateEntry]) -> list[EntryGroup]: + """Group a thematic group's entries by section, both in README order, listing each entry once.""" + placed: set[tuple[str, str]] = set() + groups: list[EntryGroup] = [] + for section in sections: + entries: list[TemplateEntry] = [] + for parsed in section["entries"]: + key = (parsed["url"], parsed["name"]) + if key not in placed: + placed.add(key) + entries.append(entries_by_key[key]) + groups.append(EntryGroup(name=section["name"], slug=section["slug"], url=category_path(section), entries=entries)) + return groups + + def build_breadcrumb_json_ld(items: Sequence[tuple[str, str]]) -> dict: return { "@type": "BreadcrumbList", @@ -404,6 +467,21 @@ def link_llms_category_index_to_canonical_pages(markdown: str, categories: Seque return "".join(out) +def link_description_anchors_to_category_pages(categories: Sequence[ParsedSection]) -> None: + """Point README anchor links in section descriptions at category pages, which lack those anchors.""" + category_paths = {} + for category in categories: + category_paths[f"#{category['slug']}"] = category_path(category) + category_paths[github_markdown_anchor(category["name"])] = category_path(category) + + def replace_attrs(match: re.Match[str]) -> str: + path = category_paths.get(match.group(1)) + return f'href="{path}"' if path else match.group(0) + + for category in categories: + category["description_html"] = ANCHOR_LINK_ATTRS_RE.sub(replace_attrs, category["description_html"]) + + def build_llms_txt( template_text: str, *, @@ -583,6 +661,7 @@ def build(repo_root: Path) -> None: duplicates = {s for s, n in Counter(all_top_level_slugs).items() if n > 1} if duplicates: raise ValueError(f"slug collision in /categories/ namespace: {sorted(duplicates)}. Rename a category or group so their slugs differ.") + link_description_anchors_to_category_pages(categories) total_entries = sum(c["entry_count"] for c in categories) entries = extract_entries(categories, parsed_groups) build_date = datetime.now(UTC) @@ -617,12 +696,7 @@ def build(repo_root: Path) -> None: filter_urls: dict[str, str] = dict(category_urls) for group in parsed_groups: filter_urls[group["name"]] = group_path(group["slug"]) - for entry in entries: - for sub in entry.get("subcategories", []): - filter_urls[sub["value"]] = sub["url"] builtin_entries = [e for e in entries if e.get("source_type") == BUILTIN_FILTER] - if builtin_entries: - filter_urls[BUILTIN_FILTER] = BUILTIN_PATH env = Environment( loader=FileSystemLoader(website / "templates"), @@ -635,7 +709,6 @@ def build(repo_root: Path) -> None: shutil.rmtree(site_dir) site_dir.mkdir(parents=True) - filter_urls_json = json.dumps(filter_urls, sort_keys=True, ensure_ascii=False).replace(" None: sponsors=sponsors, category_urls=category_urls, filter_urls=filter_urls, - filter_urls_json=filter_urls_json, homepage_json_ld=homepage_json_ld, ), encoding="utf-8", @@ -672,11 +744,13 @@ def build(repo_root: Path) -> None: page_dir: Path, parent_category: ParsedSection | None = None, group_categories: Sequence[ParsedSection] | None = None, + entry_groups: Sequence[EntryGroup] = (), ) -> None: page_dir.mkdir(parents=True, exist_ok=True) parent_name = parent_category["name"] if parent_category else None category_title = category_meta_title(category["name"], parent_name) - category_description = category_meta_description(category["name"], len(entries), category["description"], parent_name) + intro_html, guide_html, intro_lead = load_category_intro(website / "data" / "category_intros" / f"{current_path.removeprefix('/categories/').strip('/')}.md") + category_description = intro_lead or category_meta_description(category["name"], len(entries), category["description"], parent_name) breadcrumbs = [("Awesome Python", SITE_URL)] if parent_category: breadcrumbs.append((parent_category["name"], category_public_url(parent_category))) @@ -691,12 +765,14 @@ def build(repo_root: Path) -> None: category_title=category_title, category_url=category_url, category_description=category_description, + intro_html=intro_html, + guide_html=guide_html, entries=entries, + entry_groups=entry_groups, total_categories=len(categories), category_urls=category_urls, current_path=current_path, filter_urls=filter_urls, - filter_urls_json=filter_urls_json, parent_category=parent_category, group_categories=group_categories, category_json_ld=category_json_ld, @@ -704,6 +780,8 @@ def build(repo_root: Path) -> None: encoding="utf-8", ) + entries_by_key = {(e["url"], e["name"]): e for e in entries} + section_groups = {category["name"]: group_section_entries(category, entries_by_key) for category in categories} for category in categories: render_category( category, @@ -711,6 +789,7 @@ def build(repo_root: Path) -> None: entries=[e for e in entries if category["name"] in e["categories"]], current_path=category_path(category), page_dir=categories_dir / category["slug"], + entry_groups=section_groups[category["name"]], ) for group in parsed_groups: @@ -721,6 +800,7 @@ def build(repo_root: Path) -> None: current_path=group_path(group["slug"]), page_dir=categories_dir / group["slug"], group_categories=group["categories"], + entry_groups=group_entries_by_section(group["categories"], entries_by_key), ) if builtin_entries: @@ -770,8 +850,21 @@ def build(repo_root: Path) -> None: current_path=subcategory_path(cat_slug, sub_slug), page_dir=categories_dir / cat_slug / sub_slug, parent_category=cat_by_slug[cat_slug], + entry_groups=[EntryGroup(name="", slug="", url="", entries=group["entries"]) for group in section_groups[cat_by_slug[cat_slug]["name"]] if group["name"] == sub_name], ) + redirects_file = website / "data" / "redirects.json" + redirects = json.loads(redirects_file.read_text(encoding="utf-8")) if redirects_file.exists() else {} + for old_path, new_path in redirects.items(): + tpl_redirect = env.get_template("redirect.html") + stub = site_dir / old_path.strip("/") / "index.html" + if stub.exists(): + raise ValueError(f"redirect source {old_path} is a live page; remove it from redirects.json") + if not (site_dir / new_path.strip("/") / "index.html").exists(): + raise ValueError(f"redirect target {new_path} for {old_path} does not exist") + stub.parent.mkdir(parents=True, exist_ok=True) + stub.write_text(tpl_redirect.render(target_url=SITE_URL + new_path.lstrip("/")), encoding="utf-8") + static_src = website / "static" static_dst = site_dir / "static" if static_src.exists(): diff --git a/website/data/category_intros/admin-panels.md b/website/data/category_intros/admin-panels.md new file mode 100644 index 00000000..a318e6c4 --- /dev/null +++ b/website/data/category_intros/admin-panels.md @@ -0,0 +1,17 @@ +Rather than hand-build a Python admin panel, use Flask-Admin in Flask apps. Django has one built in, which Unfold restyles for dashboards and Grappelli reskins. + +How to choose: + +- A Flask app: Flask-Admin +- Dashboards and a Tailwind CSS look on the Django admin: Unfold +- A grid-based skin for the Django admin: Grappelli + +Flask-Admin [solves the boring problem](https://flask-admin.readthedocs.io/en/stable/) of building an admin interface on top of an existing data model. Where the Django admin works from sensible defaults, Flask-Admin [leaves it to you](https://flask-admin.readthedocs.io/en/stable/advanced/#migrating-from-django) to tell it what to display and how. Register [a `ModelView` for each model](https://flask-admin.readthedocs.io/en/stable/introduction/#adding-model-views) with `admin.add_view()` to get its list, create, and edit views. To shape them, subclass `ModelView` and set attributes like `column_list` and `column_searchable_list`. It works with several ORMs and MongoDB, and if you don't know where to start, its docs point you to [SQLAlchemy](https://flask-admin.readthedocs.io/en/stable/advanced/#using-different-database-backends). For a page that isn't tied to a model, like analytics, extend [`BaseView`](https://flask-admin.readthedocs.io/en/stable/introduction/#standalone-views). + +Unfold is [a modern Django admin theme](https://unfoldadmin.com/) for building dashboards, internal tools, and business applications with Tailwind CSS. Put `"unfold"` [first in `INSTALLED_APPS`](https://unfoldadmin.com/docs/installation/quickstart/), before `django.contrib.admin`. Your admin URLs stay the same. Your admin classes then inherit from `unfold.admin.ModelAdmin`, since the default class gives you unstyled forms without Unfold's features. That includes the User and Group admins Django registers for you, which you [unregister and register again](https://unfoldadmin.com/docs/installation/auth/) with Unfold's class. Build [the dashboard](https://unfoldadmin.com/docs/configuration/dashboard/) by overriding the admin index template, fed with data from a callback. + +Grappelli is [a grid-based skin](https://django-grappelli.readthedocs.io/en/latest/) for the Django admin. Its FAQ says that [if you're pleased with how the original admin looks](https://django-grappelli.readthedocs.io/en/latest/faq.html#why-should-i-use-grappelli), you probably shouldn't use it. Put `grappelli` before `django.contrib.admin` in `INSTALLED_APPS`, and [include its URLs](https://django-grappelli.readthedocs.io/en/latest/quickstart.html#setup), which related lookups and autocompletes need. Its [customization examples](https://django-grappelli.readthedocs.io/en/latest/customization.html) build on Django's own `admin.ModelAdmin` and inline classes. They add options for collapsible fieldsets, drag-and-drop inline sorting, and autocomplete lookups. + +An admin panel is for trusted users. Django's docs [limit the admin](https://docs.djangoproject.com/en/stable/ref/contrib/admin/#overview) to an organization's internal management tool, not your whole front end. By default, it lets in only users with `is_staff` set. Unfold and Grappelli both run on it, so the same holds for them. When you need a process-centric interface instead of one built around tables and fields, write your own views. Flask-Admin leaves [keeping unwanted users out](https://flask-admin.readthedocs.io/en/stable/introduction/#authorization-permissions) to you: override `is_accessible` on your admin views with your own login check. + +On Django, admin add-ons built for the stock look may not match a theme. Before you switch, check [Unfold's integrations](https://unfoldadmin.com/docs/) or [Grappelli's third-party list](https://django-grappelli.readthedocs.io/en/latest/thirdparty.html) for the ones you rely on. diff --git a/website/data/category_intros/ai-and-agents.md b/website/data/category_intros/ai-and-agents.md new file mode 100644 index 00000000..777525a0 --- /dev/null +++ b/website/data/category_intros/ai-and-agents.md @@ -0,0 +1,48 @@ +LangChain is the place to start among Python libraries for AI agents, and LangGraph gives you control of every step. vLLM serves your own models. + +How to choose: + +- Skills for your coding agent: Django AI Skills, Sentry Skills, or Trail of Bits Skills +- A first agent: LangChain, or LangGraph to control every step +- An agent built on one vendor's platform: OpenAI Agents SDK or Claude Agent SDK +- A ready-made personal assistant: Hermes Agent, or AstrBot for chat apps +- Prompts tuned against a metric instead of by hand: DSPy +- Structured output, RAG, or agent memory: Instructor, LlamaIndex, or Mem0 +- Running pre-trained models: Transformers +- Serving a model: vLLM, or MLX LM on Apple silicon +- One API for many LLM providers: LiteLLM +- Image and video generation: Diffusers +- Fine-tuning: PEFT, Unsloth, or Axolotl +- Speech: Whisper for speech to text, Kitten TTS for text to speech + +New to agents? LangGraph's own docs [recommend LangChain's prebuilt agents](https://docs.langchain.com/oss/python/langgraph/overview), which run on LangGraph: give an agent a model, tools, and a prompt, and the loop is handled for you. Drop down to LangGraph for [needs that combine deterministic and agentic workflows](https://docs.langchain.com/oss/python/langchain/overview). You don't need LangChain to use LangGraph. + +Pydantic AI is the pick when you want your type checker to cover the agent too. Give the agent an output type, and [every run comes back as a validated Pydantic model](https://pydantic.dev/docs/ai/overview/); when validation fails, the model is asked to try again. Tools and instructions get their dependencies through typed injection, so you can swap in a test double in unit tests. + +CrewAI splits the work into Crews, teams of role-playing agents, and Flows, event-driven workflows that hold state. For production apps, its docs recommend [starting with a Flow](https://docs.crewai.com/en/concepts/production-architecture) and calling Crews from it. + +OpenAI Agents SDK keeps [the primitives few](https://openai.github.io/openai-agents-python/): agents, handoffs, and guardrails, with tracing built in. It also runs [non-OpenAI models](https://openai.github.io/openai-agents-python/models/). Claude Agent SDK runs [Claude Code as a library](https://code.claude.com/docs/en/agent-sdk/overview): the same built-in tools, permissions, sessions, and hooks, inside your own process. + +Instructor gets validated data out of an LLM into a Pydantic model, with retries when validation fails. Its own docs draw the line: [Instructor for extraction, Pydantic AI for agents](https://python.useinstructor.com/). + +DSPy has you [write signatures, not prompts](https://dspy.ai/). Give it examples and a metric, and its optimizers tune the prompts for you. + +LlamaIndex is a [data framework](https://github.com/run-llama/llama_index) for LLM apps: it loads, indexes, and queries your documents. Install `llama-index` to start, or `llama-index-core` plus only the integrations you need. + +Mem0 adds [memory that persists across sessions](https://docs.mem0.ai/). Self-host the open-source version, or use the managed platform. OpenViking is [AGPL-licensed](https://github.com/volcengine/OpenViking/blob/main/LICENSE), where Mem0 is Apache-licensed. + +Transformers runs pre-trained models from the Hugging Face Hub. Start with [`pipeline()`](https://huggingface.co/docs/transformers/pipeline_tutorial): pick a task and a model, and it handles preprocessing and output. Diffusers works the same way for [diffusion models](https://huggingface.co/docs/diffusers/index). + +vLLM serves a model behind an [OpenAI-compatible API](https://docs.vllm.ai/en/latest/getting_started/quickstart.html). SGLang does too, and its [RadixAttention caches shared prefixes](https://docs.sglang.io/), which helps when requests share a long prompt. On Apple silicon, [MLX LM](https://github.com/ml-explore/mlx-lm) runs and fine-tunes models locally. + +LiteLLM puts many LLM providers behind one OpenAI-style API. Use [the Python SDK in your code, or run the proxy as a gateway](https://docs.litellm.ai/docs/) when a platform team needs keys, budgets, and spend tracking across projects. + +PEFT [trains a small set of extra parameters](https://huggingface.co/docs/peft/index) instead of the whole model, and works with Transformers and Diffusers. Unsloth's docs [recommend starting with QLoRA](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide). Axolotl drives [the whole pipeline from one YAML file](https://docs.axolotl.ai/): preprocessing, training, evaluation, quantization, and inference. + +Whisper is a [general-purpose speech recognition model](https://github.com/openai/whisper) that also translates speech and identifies languages. Microsoft marks VibeVoice for [research and development only](https://github.com/microsoft/VibeVoice). + +For text to speech, [Kitten TTS runs on CPU](https://github.com/KittenML/KittenTTS) without a GPU. gTTS calls [Google Translate's undocumented speech endpoint](https://github.com/pndurette/gTTS), so it needs the internet and can break without notice. + +The skill repos aren't pip packages: they install into your coding agent, not your app. Django AI Skills and Sentry Skills follow the [Agent Skills](https://agentskills.io/) open format. + +Write your app against the OpenAI API format, and you can switch between a hosted model and your own: vLLM, SGLang, and the LiteLLM proxy all speak it. diff --git a/website/data/category_intros/algorithms-and-design-patterns.md b/website/data/category_intros/algorithms-and-design-patterns.md new file mode 100644 index 00000000..21f6bad0 --- /dev/null +++ b/website/data/category_intros/algorithms-and-design-patterns.md @@ -0,0 +1,24 @@ +Read The Algorithms to see how algorithms work, not to ship. Install Sorted Containers when your code needs a Python algorithms library for sorted collections. + +How to choose: + +- Sorted lists, dicts, and sets: Sorted Containers +- Learning how an algorithm works: The Algorithms +- Algorithm code to read that also installs with pip: algorithms +- A state machine bound to an object you already have: transitions +- Design patterns, and which ones Python doesn't need: python-patterns +- A state machine declared as a class, up to full statecharts: python-statemachine + +Python's standard library is great [until you need a sorted collections type](https://grantjenks.com/docs/sortedcontainers/). Sorted Containers fills that gap in pure Python, with no C compiler to install. `SortedList` is its core type: it [keeps its values in ascending order](https://grantjenks.com/docs/sortedcontainers/introduction.html#sorted-list) as you add them with `add()` or `update()`. `SortedDict` is a dict that also keeps a sorted list of its keys, and `SortedSet` is a set that keeps a sorted list of its values. + +The Algorithms implements [algorithms in Python for education](https://github.com/TheAlgorithms/Python), from sorts and searches to graphs and dynamic programming. Browse them by topic in its [directory](https://github.com/TheAlgorithms/Python/blob/master/DIRECTORY.md). + +The algorithms package puts each data structure and algorithm [in a self-contained file](https://github.com/keon/algorithms) with docstrings, type hints, and complexity notes, written to be read and learned from. It also installs with pip, so your code can import what you've read, like `from algorithms.graph import dijkstra`. + +transitions is [a lightweight, object-oriented state machine](https://github.com/pytransitions/transitions) that you bind to an object you already have. [Pass `Machine`](https://github.com/pytransitions/transitions#basic-initialization) your model, its states, and its transitions as dicts, each with a trigger, a source, and a destination. The model then gets a method for each trigger, like `evaporate()`. + +python-patterns is [a collection of design patterns and idioms](https://github.com/faif/python-patterns), one file per pattern, grouped as creational, structural, behavioral, and more. Its README asks you to care more about why you pick a pattern than how you implement it. [Its anti-patterns section](https://github.com/faif/python-patterns#-anti-patterns) lists the ones not recommended in Python: modules are already singletons, so use a module-level variable instead of a Singleton class. + +python-statemachine defines [flat state machines or full statecharts](https://python-statemachine.readthedocs.io/en/latest/) in a declarative class that works in both sync and async code. Statecharts add compound states, parallel regions, and history. States are class attributes, like `green = State(initial=True)`. `green.to(yellow)` declares a transition, and `|` combines transitions into one event, as in `cycle = green.to(yellow) | yellow.to(red)`. + +The Algorithms says its implementations [may be less efficient than the standard library's](https://github.com/TheAlgorithms/Python), so where the standard library has an algorithm, use its version in your code. python-statemachine's docs show how to [rewrite a hand-written State pattern declaratively](https://python-statemachine.readthedocs.io/en/latest/how-to/coming_from_state_pattern.html), like the one in python-patterns' `state.py`. For more places to learn and practice algorithms, see [awesome-algorithms](https://github.com/tayllan/awesome-algorithms). diff --git a/website/data/category_intros/asynchronous-programming.md b/website/data/category_intros/asynchronous-programming.md new file mode 100644 index 00000000..f71a36b9 --- /dev/null +++ b/website/data/category_intros/asynchronous-programming.md @@ -0,0 +1,30 @@ +I/O-bound or CPU-bound? I/O-bound code wants a Python async library, and asyncio comes built in. CPU-bound work goes to a concurrent.futures process pool. + +How to choose: + +- I/O-bound code written with async/await: asyncio +- CPU-bound work, or blocking calls, in a pool of processes or threads: concurrent.futures +- Trio-style task groups and cancel scopes on asyncio, or a library that runs on both asyncio and Trio: AnyIO +- A faster, drop-in event loop for an asyncio app on Linux or macOS: uvloop +- Structured concurrency from the ground up, with its own libraries: Trio +- Existing synchronous code, made concurrent without async/await: gevent +- Network servers and clients with protocols built in, like SSH, mail, and DNS: Twisted +- Processes you manage yourself, talking through queues and pipes: multiprocessing + +asyncio is the standard library's way to [write concurrent code with async/await](https://docs.python.org/3/library/asyncio.html), and many async web servers, database drivers, and task queues build on it. Start your program with `asyncio.run()`, [called once as the main entry point](https://docs.python.org/3/library/asyncio-runner.html#asyncio.run). App code should [rarely need the event loop object](https://docs.python.org/3/library/asyncio-eventloop.html) itself. Run related tasks in an `asyncio.TaskGroup`: when one task fails, it [cancels the rest, which `gather()` doesn't](https://docs.python.org/3/library/asyncio-task.html#asyncio.gather). CPU-heavy code holds up every task on the loop, so [run it in another process](https://docs.python.org/3/library/asyncio-dev.html#asyncio-handle-blocking): hand it to a `ProcessPoolExecutor` with `loop.run_in_executor()`. + +concurrent.futures runs callables on threads or processes behind [the same interface](https://docs.python.org/3/library/concurrent.futures.html). `ThreadPoolExecutor` is [for overlapping I/O](https://docs.python.org/3/library/concurrent.futures.html#threadpoolexecutor). For CPU-bound work on a multi-core machine, the threading docs [advise processes](https://docs.python.org/3/library/threading.html) instead, and `ProcessPoolExecutor` runs them for you. It [takes only picklable functions and arguments](https://docs.python.org/3/library/concurrent.futures.html#processpoolexecutor), though. Use either executor [in a `with` block](https://docs.python.org/3/library/concurrent.futures.html#concurrent.futures.Executor.shutdown), which shuts it down and waits for its work to finish. + +AnyIO brings [Trio-like structured concurrency to asyncio](https://anyio.readthedocs.io/en/stable/). Code written against its API runs unmodified on asyncio or Trio, so a library built on it doesn't choose for its users. Its docs see [strong merits in its APIs for applications too](https://anyio.readthedocs.io/en/stable/why.html), starting with cancel scopes for [more predictable cancellation](https://anyio.readthedocs.io/en/stable/why.html#design-problems-with-cancellation). Start with `anyio.run(main)`, which [runs on asyncio unless you pass `backend="trio"`](https://anyio.readthedocs.io/en/stable/basics.html#running-async-programs). Spawn tasks in a [task group](https://anyio.readthedocs.io/en/stable/tasks.html): when one child task raises, the rest are cancelled. + +uvloop is [a drop-in replacement for asyncio's event loop](https://github.com/MagicStack/uvloop), built on libuv. Its README prefers `uvloop.run(main())`, which configures `asyncio.run()` to use uvloop, so the rest of your asyncio code stays the same. uvloop runs on Linux and macOS. + +Trio has [an obsessive focus on usability and correctness](https://trio.readthedocs.io/en/stable/). Child tasks run in a nursery, opened with `async with trio.open_nursery()`, and [Trio never discards their exceptions](https://trio.readthedocs.io/en/stable/tutorial.html#okay-let-s-see-something-cool-already). Functions [take no timeout arguments](https://trio.readthedocs.io/en/stable/reference-core.html#blocking-and-non-blocking-methods): you wrap the code in a cancel scope like `trio.move_on_after()`. Trio runs its own event loop, so asyncio functions [don't work inside `trio.run()`](https://trio.readthedocs.io/en/stable/tutorial.html#task-switching-illustrated). Check [the Trio library list](https://trio.readthedocs.io/en/stable/awesome-trio-libraries.html) for what you need first. + +gevent uses greenlets to give you [a synchronous API on top of an event loop](https://www.gevent.org/intro.html). Its monkey patching swaps the standard library's blocking sockets for cooperative ones, so code that knows nothing about gevent runs concurrently. Most programs should [patch everything with `monkey.patch_all()`](https://www.gevent.org/intro.html#beyond-sockets), and the main module should do it [before any other imports](https://www.gevent.org/api/gevent.monkey.html). + +Twisted is [an event-based framework for internet applications](https://github.com/twisted/twisted) that ships clients and servers for HTTP, SSH, IMAP, POP3, SMTP, DNS, IRC, and XMPP. Write new Twisted code as `async def` coroutines, which its docs [prefer over `inlineCallbacks`](https://docs.twisted.org/en/stable/core/howto/defer-intro.html#inline-callbacks-using-yield), and start one with [`Deferred.fromCoroutine()`](https://docs.twisted.org/en/stable/core/howto/defer-intro.html#coroutines-with-async-await). + +multiprocessing runs work in subprocesses, so it can [use every processor on a machine](https://docs.python.org/3/library/multiprocessing.html#introduction). Its own docs point to `ProcessPoolExecutor` as the higher-level interface for pooled tasks. Reach for multiprocessing when you need what it adds, like killing a running process or passing data through queues and pipes. Its guidelines say to [avoid shared state](https://docs.python.org/3/library/multiprocessing.html#all-start-methods) and keep the data moving between processes small. + +On an event loop, a blocking call holds up every other task, so push it to a worker thread: asyncio has [`asyncio.to_thread()`](https://docs.python.org/3/library/asyncio-task.html#asyncio.to_thread), AnyIO [`to_thread.run_sync()`](https://anyio.readthedocs.io/en/stable/threads.html), and Trio [`trio.to_thread.run_sync()`](https://trio.readthedocs.io/en/stable/reference-core.html#threads-if-you-must). When multiprocessing or `ProcessPoolExecutor` runs your code in child processes, define its functions in a module. Guard your entry point with `if __name__ == '__main__':` too, so each new process can [import your main module safely](https://docs.python.org/3/library/multiprocessing.html#multiprocessing-safe-main-import). diff --git a/website/data/category_intros/audio-video-processing.md b/website/data/category_intros/audio-video-processing.md new file mode 100644 index 00000000..1e5236b5 --- /dev/null +++ b/website/data/category_intros/audio-video-processing.md @@ -0,0 +1,27 @@ +Each job has its own pick: librosa is the Python audio processing library for music analysis, MoviePy edits video from a script, and Mutagen tags audio files. + +How to choose: + +- Cutting, joining, and fading audio files: pydub +- Music and audio analysis: librosa +- Editing video or making GIFs from a script: MoviePy +- Real-time video from cameras and network streams: VidGear +- Reading and writing tags across formats: Mutagen +- Reading tags only: tinytag +- Tagging and organizing your music collection: beets + +pydub gives you a simple, high-level interface to cut, join, and fade audio. It opens and saves WAV files in pure Python, but [other formats like MP3 need FFmpeg](https://github.com/jiaaro/pydub#dependencies), so install FFmpeg with it. An AudioSegment is [immutable](https://github.com/jiaaro/pydub#quickstart): every operation returns a new one, so you can chain them, and every length and position is in milliseconds. + +librosa gives you [the foundational algorithms and tools for music information retrieval](https://librosa.org/doc/latest/index.html). By default, `librosa.load` resamples the signal and mixes stereo down to mono, and those defaults [suit most analysis tasks](https://librosa.org/doc/latest/auto_tutorials/01-intro/01-load.html); pass `sr=None` to keep the file's own sampling rate. When you need more control than `load` gives you, such as writing files, [its docs recommend using its audio I/O backend directly](https://librosa.org/doc/latest/ioformats.html). + +MoviePy is for [automating video editing](https://zulko.github.io/moviepy/getting_started/quick_presentation.html): processing many videos, composing them in complicated ways, or making videos and GIFs on a web server. A script loads clips, modifies them, puts them together, and writes the result. Modifying a clip [returns a new clip and leaves the original alone](https://zulko.github.io/moviepy/user_guide/modifying.html), and the computation happens at the final render. Open file clips in a `with` block, or [call `close()`](https://zulko.github.io/moviepy/user_guide/loading.html) when you're done, since each one holds a subprocess and a lock on the file. MoviePy can't stream video. For frame-by-frame analysis, its docs send you to a computer vision library. + +VidGear is a framework for [real-time media applications](https://abhitronix.github.io/vidgear/latest/) built on OpenCV and FFmpeg. All its APIs [keep OpenCV's coding syntax](https://abhitronix.github.io/vidgear/latest/switch_from_cv/). Each task has [its own gear](https://abhitronix.github.io/vidgear/latest/gears/): CamGear reads cameras, network streams, and streaming sites in multiple threads. WriteGear writes frames to a video file or network stream, and StreamGear transcodes video into adaptive streaming formats. [Install OpenCV first](https://abhitronix.github.io/vidgear/latest/installation/pip_install/), since the core functions need it. + +Mutagen reads and writes tags with [roughly the same API across all tag formats](https://mutagen.readthedocs.io/en/latest/). `mutagen.File` [guesses the file type](https://mutagen.readthedocs.io/en/latest/user/gettingstarted.html). ID3 tags in MP3 files are highly structured; for common keys, use [the simpler EasyID3 interface](https://mutagen.readthedocs.io/en/latest/user/id3.html). Mutagen is GPL-licensed; if you only read tags, MIT-licensed tinytag avoids that. + +tinytag only reads metadata, and [writing support will not be added](https://github.com/tinytag/tinytag): its README points you to Mutagen for that. It's pure Python with no dependencies and gives you the same API for every format. `TinyTag.get()` returns an object with attributes like `artist` and `duration`. + +beets is a command-line music library manager, not a library you import: it [catalogs your collection and improves its metadata](https://beets.io/) as it goes. Install it [as a standalone tool](https://beets.readthedocs.io/en/stable/guides/installation.html), isolated from your system Python and other packages. `beet import` can modify and move your files, so [back up first and import a few albums at a time](https://beets.readthedocs.io/en/stable/guides/main.html). [Plugins](https://beets.readthedocs.io/en/stable/plugins/index.html) add commands, fetch extra data during import, and add metadata sources. + +FFmpeg sits under most of these projects: pydub needs it for any format other than WAV, MoviePy runs on it, and VidGear's WriteGear and StreamGear wrap it. When you only want to convert a video file or turn images into a movie, [call FFmpeg directly](https://zulko.github.io/moviepy/getting_started/quick_presentation.html). MoviePy's own docs say it's faster and uses less memory than going through MoviePy. diff --git a/website/data/category_intros/authentication.md b/website/data/category_intros/authentication.md new file mode 100644 index 00000000..a02aef1a --- /dev/null +++ b/website/data/category_intros/authentication.md @@ -0,0 +1,27 @@ +Letting users log in with Google takes one Python authentication library: django-allauth on Django, which does password sign-up too, or Authlib elsewhere. + +How to choose: + +- Sign-up, password login, and Google or GitHub login on a Django site: django-allauth +- Google or GitHub login, OAuth API calls, or an OAuth server outside Django: Authlib +- Adding OAuth to a web framework or HTTP library you maintain: oauthlib +- Your own OAuth 2.0 server on Django, or tokens for a Django REST framework API: django-oauth-toolkit +- Signing and verifying JSON Web Tokens: PyJWT +- Permissions you grant per object to users and groups: django-guardian +- Permissions that follow from the object's data, with no extra tables: django-rules + +django-allauth exists to handle [local and social accounts in one app](https://docs.allauth.org/en/latest/introduction/index.html): a user can sign up with a password or log in with Google. Follow the [quickstart](https://docs.allauth.org/en/latest/installation/quickstart.html): add its authentication backend, apps, and middleware, then include `allauth.urls` under `accounts/`. Its login, logout, and password views can replace the ones in `django.contrib.auth.urls`. Give each OAuth provider its client ID and secret in a [`SocialApp` or the `SOCIALACCOUNT_PROVIDERS` setting](https://docs.allauth.org/en/latest/installation/quickstart.html#post-installation). For a single-page or mobile app, add [`allauth.headless`](https://docs.allauth.org/en/latest/headless/introduction.html). + +Authlib covers both sides of OAuth 2.0 and OpenID Connect. As a client, it comes in [two kinds](https://docs.authlib.org/en/latest/oauth2/client/index.html): HTTP clients for scripts and service-to-service calls, and web clients that log users in through Flask, Django, Starlette, or FastAPI. For login, [create an `OAuth` registry and `register()` each provider](https://docs.authlib.org/en/latest/oauth2/client/web/index.html). For an OpenID Connect provider like Google, pass its discovery URL as [`server_metadata_url`](https://docs.authlib.org/en/latest/oauth2/client/web/index.html#parsing-id-token), and Authlib reads the other endpoints from it. To [run your own OAuth 2.0 or OpenID Connect provider](https://docs.authlib.org/en/latest/oauth2/authorization-server/index.html), it has server integrations for Flask and Django. + +PyJWT [encodes and decodes JSON Web Tokens](https://pyjwt.readthedocs.io/en/latest/). Treat every token as untrusted input: [hard-code the `algorithms` you pass to `jwt.decode()`](https://pyjwt.readthedocs.io/en/latest/algorithms.html#specifying-an-algorithm), never read them from the token's own header, and don't mix HS and RS algorithms. For tokens from an OpenID Connect provider, point [`PyJWKClient` at its JWKS endpoint](https://pyjwt.readthedocs.io/en/latest/usage.html#retrieve-rsa-signing-keys-from-a-jwks-endpoint) instead of hard-coding its public keys. + +oauthlib implements OAuth's logic [without assuming a web framework or HTTP request object](https://github.com/oauthlib/oauthlib), so a framework or library maintainer can write a thin layer on top and get OAuth support. Its FAQ says [most people use it indirectly](https://oauthlib.readthedocs.io/en/latest/faq.html#how-do-i-use-oauthlib-with-google-twitter-and-other-providers): django-oauth-toolkit and django-allauth both build on it. To build a provider on it, most of the work is [a `RequestValidator`](https://oauthlib.readthedocs.io/en/latest/oauth2/server.html#implement-a-validator) that maps its checks to your storage. + +django-oauth-toolkit is an [OAuth 2.0 authorization server for Django](https://django-oauth-toolkit.readthedocs.io/en/latest/): it issues and manages tokens from your existing project, and can also protect a Django or Django REST framework API. Django REST framework's docs [recommend it for OAuth 2.0](https://www.django-rest-framework.org/api-guide/authentication/#django-oauth-toolkit). [Install it](https://django-oauth-toolkit.readthedocs.io/en/latest/install.html) by adding `oauth2_provider` to `INSTALLED_APPS`, including its URLs under `o/`, and migrating. For an API, [set `OAuth2Authentication`](https://django-oauth-toolkit.readthedocs.io/en/latest/rest-framework/getting_started.html) as Django REST framework's authentication class and guard views with `TokenHasScope`. + +Django has [a foundation for object permissions but no implementation](https://docs.djangoproject.com/en/stable/topics/auth/customizing/#handling-object-permissions), and django-guardian fills it with an [extra authentication backend](https://django-guardian.readthedocs.io/en/latest/configuration/). Grant a permission on one object with `assign_perm()`, then check it with Django's own [`user.has_perm(perm, obj)`](https://django-guardian.readthedocs.io/en/latest/userguide/checks/). Django REST framework's `DjangoObjectPermissions` [names it as an example backend](https://www.django-rest-framework.org/api-guide/permissions/#djangoobjectpermissions). + +django-rules gives Django object permissions [without a database](https://github.com/dfunckt/django-rules). A permission is a rule built from predicates, plain functions like `is_book_author(user, book)` that you combine with `|`, `&`, and `~`. Add its [backend to `AUTHENTICATION_BACKENDS`](https://github.com/dfunckt/django-rules#configuring-django), register each permission with `rules.add_perm()`, and check it with `user.has_perm()`. Keep predicates and rules in their own modules, and let [`AutodiscoverRulesConfig` import each app's `rules.py`](https://github.com/dfunckt/django-rules#best-practices) at startup. + +On either side of OAuth, use PKCE. django-oauth-toolkit's security guide maps [the OAuth security recommendations of RFC 9700](https://django-oauth-toolkit.readthedocs.io/en/latest/security.html) to its settings: PKCE is one, and the implicit and password grants must not be used. Authlib's client [turns PKCE on with `code_challenge_method`](https://docs.authlib.org/en/latest/oauth2/client/web/index.html#oauth-2-0-code-challenge). Also, don't pick JWTs for sessions just because they sound stateless. django-allauth's docs point out that [if a token must stop working at logout, you need state to revoke it](https://docs.allauth.org/en/latest/headless/token-strategies/jwt-tokens.html). JWTs help most when each service checks tokens without asking a central server. diff --git a/website/data/category_intros/build-tools.md b/website/data/category_intros/build-tools.md new file mode 100644 index 00000000..376a0fcb --- /dev/null +++ b/website/data/category_intros/build-tools.md @@ -0,0 +1,18 @@ +Where Make would go, a Python build tool takes over, and SCons compiles C and C++. Invoke runs shell commands as tasks, and doit reruns only what changed. + +How to choose: + +- C, C++, or Fortran code compiled from source: SCons +- Shell commands run as named tasks with CLI flags: Invoke +- Tasks that rerun only when their input files change: doit +- Your own CLI program built from tasks: Invoke + +SCons builds software from source code. Its site calls it [an improved, cross-platform substitute for the classic Make utility](https://scons.org/), with dependency analysis for C, C++, and Fortran built in. Your build file [is a Python script](https://scons.org/doc/production/HTML/scons-user/ch02s05.html#sect-sconstruct-python) named `SConstruct`: [put `Program('hello.c')` in it](https://scons.org/doc/production/HTML/scons-user/ch02.html) and run `scons`. You declare what to build, and SCons works out the build order, [whatever order you call the builders in](https://scons.org/doc/production/HTML/scons-user/ch02s05.html#sect-order-independent). For a source tree with subdirectories, [split the build into `SConscript` files](https://scons.org/doc/production/HTML/scons-user/ch14.html#sect-sconscript-files) that the top-level `SConstruct` pulls in. SCons [puts a correct build first](https://scons.org/doc/production/HTML/scons-user/pr01.html#sect-principles): by default, it [decides a file has changed from a hash of its contents](https://scons.org/doc/production/HTML/scons-user/ch06.html#sect-contentsigs), not its modification time. + +Invoke turns your project's shell commands into Python tasks. [Write them in a `tasks.py`](https://docs.pyinvoke.org/en/stable/getting-started.html#defining-and-running-task-functions) as `@task` functions whose first argument is a context, and [run commands with `c.run()`](https://docs.pyinvoke.org/en/stable/getting-started.html#running-shell-commands). [Each parameter becomes a CLI flag](https://docs.pyinvoke.org/en/stable/getting-started.html#task-parameters), and `invoke --list` shows your tasks. A task can name [pre-tasks](https://docs.pyinvoke.org/en/stable/getting-started.html#declaring-pre-tasks) to run first, like `clean` before `build`. When one flat list of tasks gets crowded, [group them into namespaces](https://docs.pyinvoke.org/en/stable/getting-started.html#creating-namespaces) with `Collection`. + +Invoke can also [power your own CLI program](https://docs.pyinvoke.org/en/stable/concepts/library.html#reusing-invoke-s-cli-module-as-a-distinct-binary), with your tasks as its commands. It [sticks to local commands](https://www.pyinvoke.org/faq.html#why-was-invoke-split-off-from-the-fabric-project) and leaves servers and network commands to a separate library built on it. + +doit runs only what changed. [Tasks live in a `dodo.py`](https://pydoit.org/tasks.html#intro): each function whose name starts with `task_` returns a dict, and its [actions](https://pydoit.org/tasks.html#actions) are shell commands or Python functions. [List a task's `file_dep` and `targets`](https://pydoit.org/tasks.html#dependencies-targets), and doit skips it when its dependencies haven't changed and its targets exist. Inputs [don't have to be files](https://pydoit.org/dependencies.html#uptodate): an `uptodate` check such as `config_changed` reruns a task when a config string or dict changes. + +Looking for the tool that builds your package's wheels and publishes them to PyPI? That's packaging, covered in [Package Management](/categories/package-management/). diff --git a/website/data/category_intros/built-in-classes-enhancement.md b/website/data/category_intros/built-in-classes-enhancement.md new file mode 100644 index 00000000..dfe74eed --- /dev/null +++ b/website/data/category_intros/built-in-classes-enhancement.md @@ -0,0 +1,16 @@ +Skip hand-written dunder methods, and when you compare Python dataclass alternatives, pick attrs for its validators and converters. Box gives dicts dot access. + +How to choose: + +- Classes with generated dunder methods, validators, and converters: attrs +- Dicts read with dot notation, nested ones included: python-box +- Looking up a key by its value, with both directions in sync: bidict +- A faster drop-in for the `uuid` module: uuid-utils + +attrs [writes the dunder methods](https://www.attrs.org/en/latest/) that implement object protocols, so you don't have to. For new code, its docs [recommend the modern API](https://www.attrs.org/en/latest/names.html#tl-dr): `@define` on the class and `field()` for an attribute, or `@frozen` for immutable instances, with slots on by default. Type annotations stay optional. The standard library's dataclasses [gave up features for simplicity](https://www.attrs.org/en/latest/why.html#data-classes), like validators and converters, so a class that needs them goes to attrs. It [isn't a full serialization library](https://www.attrs.org/en/latest/overview.html#what-attrs-is-not), though: to serialize and validate your attrs classes, its docs point you to its sibling project, cattrs. + +python-box installs Box, a [`dict` subclass](https://github.com/cdgriffith/Box) whose keys you can also read as attributes, even ones like `"imdb stars"`. Nested dicts and lists become Box and BoxList objects, so dot access works all the way down: `movie_box.Robin_Hood_Men_in_Tights.imdb_stars`. It [reads and writes JSON, YAML, TOML, and msgpack](https://github.com/cdgriffith/Box/wiki/Converters) with methods like `from_json()` and `to_json()`, and `to_dict()` gives you plain dicts back. Keep Box for data whose keys change: attrs' docs say a dict with a fixed and known set of keys [is an object, not a hash](https://www.attrs.org/en/latest/why.html#dicts), and belongs in a class. + +bidict gives you [bidirectional mappings](https://bidict.readthedocs.io/intro.html) that work like dicts: look up a value by its key as usual, or a key by its value through `.inverse`, which stays in sync as you update the mapping. A single dict holding both directions mixes keys with values. Modeling the mapping correctly takes two one-way mappings kept in sync, [which is what bidict does](https://bidict.readthedocs.io/intro.html#why-can-t-i-just-use-a-dict) under the hood. + +uuid-utils is a [fast, drop-in replacement for Python's `uuid` module](https://aminalaee.github.io/uuid-utils/latest/), powered by Rust. Import it under the module's name, `import uuid_utils as uuid`, and call `uuid.uuid4()` or `uuid.uuid7()` as usual. Django and some other frameworks require the standard library's own `uuid.UUID` instances. For those, [import `uuid_utils.compat`](https://aminalaee.github.io/uuid-utils/latest/#compatibility-with-python-uuid) instead, which returns them and still outperforms the standard library. diff --git a/website/data/category_intros/caching.md b/website/data/category_intros/caching.md new file mode 100644 index 00000000..5cc943f9 --- /dev/null +++ b/website/data/category_intros/caching.md @@ -0,0 +1,23 @@ +In memory, on disk, or on a cache server: the Python caching library you want is cachetools, DiskCache, or dogpile.cache. For HTTP responses, it's Hishel. + +How to choose: + +- Function results in one process's memory, with a size limit or a time-to-live: cachetools +- A cache on local disk that processes on one machine share, with no server to run: DiskCache +- A Django cache backend on local disk: DiskCache +- One cache on memcached or Redis for many processes or servers: dogpile.cache +- HTTP responses in HTTPX or Requests: Hishel +- `Cache-Control` headers and response caching in FastAPI or another ASGI app: Hishel +- Django querysets, invalidated when a model changes: Cacheops + +cachetools offers [variants of the standard library's `@lru_cache`](https://cachetools.readthedocs.io/en/stable/) with more cache algorithms, including a `TTLCache` whose items [expire after a time-to-live](https://cachetools.readthedocs.io/en/stable/#cachetools.TTLCache). Wrap a function in [`@cached`](https://cachetools.readthedocs.io/en/stable/#cachetools.cached) and pass it the cache to use. The cache classes [aren't thread-safe](https://cachetools.readthedocs.io/en/stable/#cache-implementations), so when threads share a cache, give `@cached` a `threading.Lock`. The cache lives in your process's memory, so each worker process of a web app keeps its own copy. + +DiskCache is a [disk and file backed cache library](https://grantjenks.com/docs/diskcache/) in pure Python, built on SQLite, with no other process to run. Create a `Cache` with a directory path: [two `Cache` objects on the same directory](https://grantjenks.com/docs/diskcache/tutorial.html#cache) can live in separate processes, so the workers on one machine share one cache. Wrap a function in [`@cache.memoize()`](https://grantjenks.com/docs/diskcache/tutorial.html#fanoutcache), which takes arguments like `lru_cache`'s. In Django, set [`diskcache.DjangoCache`](https://grantjenks.com/docs/diskcache/tutorial.html#djangocache) as the cache `BACKEND`. Keep the directory on a local disk, since SQLite [isn't recommended on NFS mounts](https://grantjenks.com/docs/diskcache/tutorial.html#caveats). + +dogpile.cache is [a caching API over backends of any variety](https://dogpilecache.sqlalchemy.org/en/latest/), memcached and Redis among them. You ask it for a value and hand it a function that [creates the value only when needed](https://dogpilecache.sqlalchemy.org/en/latest/usage.html#overview). When the value expires, one worker regenerates it instead of every worker at once. Create a region with `make_region()` at import time, decorate functions with `@region.cache_on_arguments()`, and [call `configure()` later](https://dogpilecache.sqlalchemy.org/en/latest/usage.html#rudimentary-usage) with the backend and expiration time from your config file. When several processes share one Redis or memcached server, turn on the backend's [`distributed_lock`](https://dogpilecache.sqlalchemy.org/en/latest/api.html#dogpile.cache.backends.redis.RedisBackend) so that lock covers them all. + +Hishel caches HTTP responses by [RFC 9111](https://hishel.com/overview.html), the HTTP caching rules browsers follow, in both sync and async code. With HTTPX, [swap your client for Hishel's cache-enabled one](https://hishel.com/httpx.html#quick-start), or [give a client you already have its cache transport](https://hishel.com/httpx.html#cache-transports). With Requests, [mount its cache adapter](https://hishel.com/requests.html#quick-start) on a `Session`. In FastAPI, [a dependency sets `Cache-Control` headers](https://hishel.com/fastapi.html#quick-start) for browsers and CDNs, and its ASGI middleware also caches responses on your server. + +Cacheops caches Django querysets in Redis and [invalidates them](https://github.com/Suor/django-cacheops#user-content-invalidation) on `save()`, `delete()`, and many-to-many changes. [Add it to `INSTALLED_APPS`](https://github.com/Suor/django-cacheops#user-content-setup) and point it at its own Redis database, as its docs highly recommend. Turn caching on per app with `'app_name.*'`, since `'*.*'` can also cache tables you don't mean to, like migrations. Call `.cache()` on a queryset to cache it by hand, and wrap a function in [`@cached_as(Article)`](https://github.com/Suor/django-cacheops#user-content-usage) to drop its result whenever an `Article` changes. + +Plan how stale entries leave the cache, too. cachetools' decorators add a [`cache_clear()`](https://cachetools.readthedocs.io/en/stable/#memoizing-decorators) function, DiskCache [evicts every key with a given tag](https://grantjenks.com/docs/diskcache/tutorial.html#cache), and a function decorated by dogpile.cache takes [`invalidate()`](https://dogpilecache.sqlalchemy.org/en/latest/api.html#dogpile.cache.region.CacheRegion.cache_on_arguments) with the same arguments you'd call it with. diff --git a/website/data/category_intros/cli-development.md b/website/data/category_intros/cli-development.md new file mode 100644 index 00000000..9c2009bc --- /dev/null +++ b/website/data/category_intros/cli-development.md @@ -0,0 +1,38 @@ +Python's own argparse covers basic command-line apps. Beyond that, pick Click or Typer as your Python CLI library, and style the output with Rich. + +How to choose: + +- Basic command-line app with no dependencies: argparse +- Nested commands, or subcommands loaded lazily: Click +- Options declared once as type-hinted function parameters: Typer +- Interactive prompts and REPLs: prompt_toolkit +- CLI from existing code without writing a parser: Python Fire +- Progress bar for a loop: tqdm, or alive-progress for animated bars +- Colored text, tables, and logs: Rich, or colorama for ANSI colors on Windows +- Full-screen app in the terminal or a browser: Textual +- Terminal widgets on an event loop you already run: Urwid +- Full-screen forms and ASCII animations: asciimatics + +argparse is in the standard library. Python's docs call it [the recommended choice](https://docs.python.org/3/library/optparse.html#choosing-an-argument-parser) when you have no more specific needs, since it gives you the most out of the box for the least code. It [writes the help and usage messages](https://docs.python.org/3/library/argparse.html) and reports invalid arguments for you. For commands written as decorated functions, the same docs point to Click, and for a CLI that works with static type checking, to Typer. + +Click builds a CLI from [commands declared with decorators](https://click.palletsprojects.com/en/stable/quickstart/). You can nest them to any depth, and Click can [load subcommands lazily](https://click.palletsprojects.com/en/stable/) at runtime. It also [reads option values from environment variables](https://click.palletsprojects.com/en/stable/why/) and comes with helpers for ANSI colors, terminal size, and launching editors. + +Typer builds a CLI from your function signatures: you [declare each argument and option once](https://typer.tiangolo.com/), as a type-hinted function parameter. Write those hints with `Annotated`, which its tutorial [recommends wherever possible](https://typer.tiangolo.com/tutorial/arguments/optional/). Its docs say you get [the best results by pairing it with Rich](https://typer.tiangolo.com/tutorial/printing/): Typer structures the commands and options, and Rich displays the output. + +prompt_toolkit is for interactive input. It can [replace GNU readline](https://python-prompt-toolkit.readthedocs.io/en/stable/) or build full-screen apps. For a REPL, use a [`PromptSession`](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/asking_for_input.html), which keeps the history for the whole session. + +Python Fire turns any Python object into a CLI: [call `fire.Fire()`](https://github.com/google/python-fire/blob/master/docs/guide.md) at the end of your program. Its README pitches it for [developing and debugging your code](https://github.com/google/python-fire), exploring existing code, and turning other people's code into a CLI. It takes each argument's type from the value you pass, not from the function signature. + +tqdm adds a progress bar to any loop: [wrap the iterable in `tqdm()`](https://github.com/tqdm/tqdm). Import it from `tqdm.auto`, which picks the console bar or the Jupyter widget for you. alive-progress wraps the loop in an [`alive_bar` context manager](https://github.com/rsalmei/alive-progress) instead, with a spinner that speeds up and slows down with your throughput. + +Rich writes colored text, tables, Markdown, and syntax-highlighted code to the terminal. Start with its [drop-in `print`](https://rich.readthedocs.io/en/stable/introduction.html), then create [one `Console` at module level](https://rich.readthedocs.io/en/stable/console.html) for the rest of your app. For colored logs, send the logging module's output through [`RichHandler`](https://rich.readthedocs.io/en/stable/logging.html). + +colorama makes ANSI escape codes work on Windows and does nothing on other platforms. If that's all you need, call [`just_fix_windows_console()`](https://github.com/tartley/colorama). + +Textual apps run in the terminal, and the `textual serve` command from textual-dev [serves them in a browser](https://textual.textualize.io/guide/devtools/). Build one by [subclassing `App`](https://textual.textualize.io/guide/app/), and style its widgets with [CSS](https://textual.textualize.io/guide/CSS/). + +Urwid is a [console widget construction set](https://urwid.org/manual/overview.html) rather than a finished UI library. It runs on [your choice of event loop](https://github.com/urwid/urwid): asyncio, or another one you already use. It's [licensed under the LGPL](https://github.com/urwid/urwid/blob/master/COPYING), while Textual and asciimatics use permissive licenses. + +asciimatics does [full-screen text UIs, from interactive forms to ASCII animations](https://github.com/peterbrittain/asciimatics): create a Screen, build a Scene from Effect objects, and let the Screen play it. + +Ship your CLI as an installable package with a `[project.scripts]` entry point, not a file users run with `python`. [Click recommends it](https://click.palletsprojects.com/en/stable/entry-points/), and shell completion in Click and [Typer](https://typer.tiangolo.com/tutorial/package/) works through that entry point. Test your commands in-process: Click and Typer both provide a [`CliRunner`](https://click.palletsprojects.com/en/stable/testing/), and Textual's [`run_test()`](https://textual.textualize.io/guide/testing/) drives your app as if you were using the keyboard and mouse. diff --git a/website/data/category_intros/cli-tools.md b/website/data/category_intros/cli-tools.md new file mode 100644 index 00000000..e630e5b2 --- /dev/null +++ b/website/data/category_intros/cli-tools.md @@ -0,0 +1,26 @@ +Swap cURL for HTTPie, and your database's own client for a Python CLI tool with autocompletion: pgcli, mycli, litecli, or IRedis. + +How to choose: + +- Calling HTTP APIs from the terminal: HTTPie +- A database shell with autocompletion: pgcli for PostgreSQL, mycli for MySQL, litecli for SQLite, IRedis for Redis +- Downloading video or audio from YouTube and other sites: yt-dlp +- A new project from a template: Cookiecutter, or Copier to pull in later template changes +- A shell you script in Python: xonsh +- A tmux session with its windows and panes, from one YAML file: tmuxp + +HTTPie is a command-line HTTP client for [testing, debugging, and generally interacting with APIs](https://httpie.io/docs/cli) and HTTP servers. It formats and colors the output. A request reads like `http PUT pie.dev/put X-API-Token:123 name=John`: `Header:Value` items set headers, and `field=value` items fill the body, which HTTPie [sends as JSON by default](https://httpie.io/docs/cli/json). + +pgcli, mycli, litecli, and IRedis are terminal clients with autocompletion and syntax highlighting, one per database. Give [pgcli](https://github.com/dbcli/pgcli) a database name or a `postgresql://` URI, and [litecli](https://github.com/dbcli/litecli) the path to a SQLite file. mycli also [works with MariaDB](https://github.com/dbcli/mycli), Percona, TiDB, and Apache Doris. IRedis behaves like Redis's own client in most cases, but it's [safer on production servers](https://github.com/laixintao/iredis): it stops you from accidentally running dangerous commands like `KEYS *`. + +yt-dlp downloads audio and video from [thousands of sites](https://github.com/yt-dlp/yt-dlp). Install ffmpeg too: yt-dlp [needs it to merge separate video and audio files](https://github.com/yt-dlp/yt-dlp#dependencies). + +Cookiecutter creates projects from templates, and its templates [work for any language](https://github.com/cookiecutter/cookiecutter), not only Python. [Run `cookiecutter gh:audreyfeldroy/cookiecutter-pypackage`](https://cookiecutter.readthedocs.io/en/stable/usage.html), answer its prompts, and you get a new project from that GitHub template. + +Copier also generates projects from templates, but it calls itself [a code lifecycle management tool](https://copier.readthedocs.io/en/stable/comparisons/). When the template changes, `copier update` brings the changes into projects you already generated. [It works best](https://copier.readthedocs.io/en/stable/updating/) when both are in Git: the template tagged, your project clean, with the `.copier-answers.yml` file that records your answers. Templates can run code: Cookiecutter's [pre- and post-generate scripts](https://github.com/cookiecutter/cookiecutter), Copier's tasks. [Generate projects only from templates you trust](https://copier.readthedocs.io/en/stable/generating/). + +xonsh is a shell whose language is [a superset of Python](https://github.com/xonsh/xonsh), with shell commands built in. It isn't POSIX-compatible, so don't make it your login shell with `chsh`. Its docs [recommend a xonsh profile in your terminal emulator](https://xon.sh/install.html#before-installing) instead. + +tmuxp launches a whole tmux session [from one YAML or JSON file](https://tmuxp.git-pull.com/quickstart/): its windows, its panes, and the commands in them. Save the file as `.tmuxp.yaml` in a project, and [`tmuxp load path/to/project/`](https://github.com/tmux-python/tmuxp) builds the session. + +Several of these tools' docs have you [install them as CLI tools](https://github.com/cookiecutter/cookiecutter), and add them to a project only to use them from Python. Looking for a library to build your own Python CLI tool? That's [CLI Development](/categories/cli-development/). diff --git a/website/data/category_intros/cms.md b/website/data/category_intros/cms.md new file mode 100644 index 00000000..fa32524f --- /dev/null +++ b/website/data/category_intros/cms.md @@ -0,0 +1,12 @@ +Both Python CMS picks run on Django: Wagtail for page types your developers define, like articles, and django CMS for editors building pages on the live site. + +How to choose: + +- Content types your developers define, like articles and events: Wagtail +- Editors composing pages from reusable components on the live site: django CMS + +Wagtail is [not an instant website in a box](https://docs.wagtail.org/en/stable/getting_started/the_zen_of_wagtail.html#wagtail-is-not-an-instant-website-in-a-box): expect to write code. Start a project with [`wagtail start`](https://docs.wagtail.org/en/stable/getting_started/quick_install.html). Each page type is [a Django model that inherits from `Page`](https://docs.wagtail.org/en/stable/topics/pages.html), so give each kind of content its own type with its own fields. An event page with a date and a location [can show up in a calendar](https://docs.wagtail.org/en/stable/getting_started/the_zen_of_wagtail.html#a-cms-should-get-information-out-of-an-editor-s-head-and-into-a-database-as-efficiently-and-directly-as-possible), and a styled heading on a generic page can't. Use [StreamField](https://docs.wagtail.org/en/stable/topics/streamfield.html) for pages without a fixed structure, like blog posts, and [snippets](https://docs.wagtail.org/en/stable/topics/snippets/index.html) for content that doesn't need its own page, like headers and footers. For a headless site, use its [built-in API](https://docs.wagtail.org/en/stable/advanced_topics/api/index.html). + +In django CMS, [editors compose pages from plugins](https://docs.django-cms.org/en/stable/explanation/philosophy.html#three-disciplines-three-surfaces) in a toolbar on the live site, and developers write standard Django. Each template [declares its placeholders](https://docs.django-cms.org/en/stable/tutorials/02-templates-placeholders.html) with `{% placeholder %}`, and `CMS_TEMPLATES` lists the templates editors can pick from. For a region that is the same on every page, like a footer, use [`{% static_alias %}`](https://docs.django-cms.org/en/stable/tutorials/02-templates-placeholders.html#a-reusable-region-with-static-alias), so the content is stored once. For content from your own app, ask [where it lives](https://docs.django-cms.org/en/stable/explanation/composition.html). If it fits in a placeholder on a page, write a plugin. If it has its own list view, detail view, and URL, mount it as an app with an apphook. The core [publishes as you edit](https://docs.django-cms.org/en/stable/explanation/publishing.html), which is rarely enough for a production site with editors, so add a versioning package. + +Both are built on Django. Wagtail [deploys like a Django site](https://docs.wagtail.org/en/stable/deployment/index.html). You can add either one to an existing Django project: Wagtail [integrates into one](https://docs.wagtail.org/en/stable/getting_started/integrating_into_django.html), and django CMS [doesn't make you rebuild around it](https://docs.django-cms.org/en/stable/explanation/philosophy.html#implications-for-projects). Neither fits a very small or static site. For a two-page brochure, django CMS calls itself [overkill](https://docs.django-cms.org/en/stable/explanation/philosophy.html#when-django-cms-may-not-be-the-right-fit), and a static site generator is lighter. diff --git a/website/data/category_intros/code-analysis.md b/website/data/category_intros/code-analysis.md new file mode 100644 index 00000000..4da863c7 --- /dev/null +++ b/website/data/category_intros/code-analysis.md @@ -0,0 +1,43 @@ +Run Ruff for Python code analysis: it lints and formats. Add a type checker, since Ruff doesn't check types, and pre-commit to run Ruff on every commit. + +How to choose: + +- Rules on which modules may import which: Import Linter +- Dead code: Vulture, or repowise to index the repo for your agent +- Deeply nested functions that are hard to read: complexipy +- Several analysis tools behind one command: Prospector +- Checks before every commit: pre-commit +- Linting, formatting, and import sorting: Ruff +- Formatting without Ruff: Black, with isort on its black profile +- Deeper inference and checks of your own: Pylint +- Flake8 plugins Ruff doesn't have: Flake8 +- Security issues: Bandit +- Refactoring: Rope +- Type checking: mypy, or Pyright, ty, or Pyrefly to also check unannotated code +- Type hints from the types your code sees at runtime: MonkeyType + +Ruff lints and formats in one tool, and its FAQ lists what it [can replace](https://docs.astral.sh/ruff/faq/#which-tools-does-ruff-replace): Flake8 and dozens of its plugins, Black, and isort. To sort imports and format, [run the linter, then the formatter](https://docs.astral.sh/ruff/formatter/#sorting-imports): `ruff check --select I --fix`, then `ruff format`. When you turn on a new rule in an existing codebase, [`--add-noqa`](https://docs.astral.sh/ruff/tutorial/#adding-rules) marks the current violations, so the rule only applies to new code. + +Pick one formatter and stay with it. Ruff's formatter is designed as a drop-in replacement for Black, but it's [not meant to be used interchangeably with Black](https://docs.astral.sh/ruff/formatter/) over time. If you pick Black, set isort's black profile [in a config file at the root of your repo](https://isort.readthedocs.io/en/latest/configuration/black_compatibility.html), so it applies however isort runs. To adopt Black, [reformat everything in one commit](https://black.readthedocs.io/en/stable/guides/introducing_black_to_your_project.html) and list that commit in `.git-blame-ignore-revs`, so `git blame` skips it. + +Ruff can be a [drop-in replacement for Flake8](https://docs.astral.sh/ruff/faq/#how-does-ruffs-linter-compare-to-flake8) when your code is formatted with Black and uses few or no Flake8 plugins. Keep Flake8 when you depend on a plugin Ruff doesn't cover. In pre-commit, install the plugin [through `additional_dependencies`](https://flake8.pycqa.org/en/latest/user/using-hooks.html). + +Pylint [doesn't trust your type hints](https://pylint.readthedocs.io/en/stable/). It infers the actual values instead, which makes it slower but finds more issues in code that isn't fully typed. You can also write plugins for checks of your own. On a legacy project, start with `--errors-only`, then turn more messages on over time. Because of its speed, Pylint's docs suggest running it [in CI or a pre-push hook](https://pylint.readthedocs.io/en/stable/user_guide/installation/pre-commit-integration.html), not on every commit. + +Bandit [finds common security issues](https://github.com/PyCQA/bandit) by walking each file's syntax tree. On an existing project, [save a baseline](https://bandit.readthedocs.io/en/latest/start.html) to ignore the findings you've judged to be non-issues. When you skip a line, write `# nosec B602` instead of a bare `# nosec`, so [a new issue on that line still shows up](https://bandit.readthedocs.io/en/latest/config.html). + +Ruff is [a linter, not a type checker](https://docs.astral.sh/ruff/faq/#how-does-ruff-compare-to-mypy-or-pyright-or-pyre), so add one. mypy is [designed for gradual typing](https://mypy.readthedocs.io/en/stable/), and by default it skips functions without annotations. On an existing codebase, its docs suggest [starting with part of the code](https://mypy.readthedocs.io/en/stable/existing_code.html), running it in CI early, turning on `check_untyped_defs` as soon as you can, and aiming for `mypy --strict`. + +Pyright, ty, and Pyrefly each come with a language server for your editor, and they check unannotated code too. Pyright [checks all code regardless of annotations](https://github.com/microsoft/pyright/blob/main/docs/mypy-comparison.md). Commit its config, [run it in CI](https://github.com/microsoft/pyright/blob/main/docs/getting-started.md), and turn on strict mode file by file with `# pyright: strict`. + +ty is [designed for adoption](https://docs.astral.sh/ty/), with support for partially typed code. Add it as a [dev dependency](https://docs.astral.sh/ty/installation/), so everyone runs the same version. Pick Pyrefly when your code uses [Pydantic, Django, or attrs](https://pyrefly.org/en/docs/compare/) and you don't want the overhead of plugins. To switch to it, `pyrefly init` [migrates your existing type checker config](https://pyrefly.org/en/docs/installation/), and `pyrefly suppress` marks the current errors as ignored, so you start from a clean check. + +Import Linter checks [contracts on your imports](https://import-linter.readthedocs.io/en/stable/), such as layers: higher layers may import lower ones, not the other way around. Vulture finds unused code. For its false positives, it recommends [a whitelist over `noqa` comments](https://github.com/jendrikseipp/vulture). After you delete dead code, run it again, since it may find more. complexipy measures how hard code is for people to understand, and its docs say to [run it alongside Ruff](https://complexipy.com/comparison-with-ruff/): Ruff catches wide functions, complexipy catches deep ones. On a large existing codebase, [take a snapshot](https://complexipy.com/usage-guide/) first, so only new complexity fails. + +Prospector wraps several analysis tools and aims to be [useful out of the box](https://github.com/prospector-dev/prospector). repowise indexes your repo for coding agents and developers, with dead code and git history in the same index. It's licensed under the [AGPL, or a commercial license](https://github.com/repowise-dev/repowise). + +Rope is a refactoring library, and its wiki suggests [starting with its language server plugin](https://github.com/python-rope/rope/wiki/How-to-use-Rope-in-my-IDE-or-Text-editor%3F) before the native editor integrations. MonkeyType records the types your code sees at runtime and writes annotations from them, which its docs call [an informative first draft](https://monkeytype.readthedocs.io/en/latest/) for you to check and fix. mypy's docs suggest [collecting types from test runs](https://mypy.readthedocs.io/en/stable/existing_code.html) this way. + +pre-commit runs your hooks [before every commit](https://pre-commit.com/). Run `pre-commit install` after every clone, `pre-commit run --all-files` when you add a hook, and the same command in CI. With Ruff's hooks, [put the lint hook with `--fix` before the formatter](https://docs.astral.sh/ruff/integrations/#pre-commit), since its fixes can leave code that needs reformatting. + +The checkers here do static code analysis: they check your code without running it. MonkeyType is the exception, since it records types while your code runs. Run Ruff [in your editor](https://docs.astral.sh/ruff/editors/) and on every commit, and Pylint and your type checker in CI. diff --git a/website/data/category_intros/computer-vision.md b/website/data/category_intros/computer-vision.md new file mode 100644 index 00000000..26e1e35b --- /dev/null +++ b/website/data/category_intros/computer-vision.md @@ -0,0 +1,24 @@ +Start with OpenCV when you need a Python computer vision library for images and video. To train and run detection models, use Ultralytics YOLO. + +How to choose: + +- Image and video processing: OpenCV +- Detection, segmentation, and pose models: Ultralytics YOLO +- Vision ops inside a PyTorch model: Kornia +- Dataset curation and model evaluation: FiftyOne +- OCR on clean, printed documents: pytesseract +- OCR on text in photos: EasyOCR + +OpenCV comes as [four pip packages that share the `cv2` namespace](https://github.com/opencv/opencv-python), so install only one: opencv-python for the main modules, or opencv-contrib-python to add the extra modules. If you never call `cv2.imshow` or you build your GUI with another toolkit, install the headless variant of either one, which also makes Docker images smaller. + +Ultralytics YOLO covers the [whole life of a model](https://docs.ultralytics.com/modes/): train, validate, predict, export, and track, from Python or the `yolo` command. Its docs recommend [starting training from a pretrained model](https://docs.ultralytics.com/modes/train/). To deploy, [export it](https://docs.ultralytics.com/modes/export/) to ONNX, TensorRT, CoreML, or another format. The code and the models you train with it are [AGPL-3.0](https://www.ultralytics.com/license), so unless you open-source your whole project, you need an Enterprise License. + +Kornia is a [differentiable computer vision library like OpenCV, with strong GPU support](https://kornia.readthedocs.io/en/latest/get-started/introduction.html). Every operator works on PyTorch tensors and supports autograd, so vision ops can run on the GPU and sit inside your training loop. + +FiftyOne works on the data side of a model. [Load your dataset and your model's predictions into it](https://docs.voxel51.com/user_guide/basics.html), then see where the model succeeds and fails, and find mistakes in your labels. It [integrates with Ultralytics](https://docs.voxel51.com/integrations/ultralytics.html), so you can run and fine-tune YOLO models on FiftyOne datasets. + +pytesseract [wraps the Tesseract OCR engine](https://github.com/madmaze/pytesseract), which you install on its own, then put on your PATH or point `tesseract_cmd` at. Tesseract suits clean, printed text and needs no GPU. To get better results, [improve the image first](https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html): Tesseract works best at 300 DPI or more, and retraining rarely helps unless you use an unusual font or a new language. + +EasyOCR is a [general OCR that reads text in photos as well as in documents](https://www.jaided.ai/easyocr), in dozens of languages. It handles text in photos, where Tesseract struggles, but it's slow without a GPU. Create a `Reader` for your languages [once](https://github.com/JaidedAI/EasyOCR) and reuse it for every image, since that call loads the model into memory. + +Ultralytics YOLO, Kornia, and EasyOCR run on PyTorch. When you need a specific CUDA build, install PyTorch before the library, as [Ultralytics](https://docs.ultralytics.com/quickstart/) and [Kornia](https://kornia.readthedocs.io/en/latest/get-started/installation.html) recommend. EasyOCR's README says the same for Windows. diff --git a/website/data/category_intros/configuration-files.md b/website/data/category_intros/configuration-files.md new file mode 100644 index 00000000..e6e2921a --- /dev/null +++ b/website/data/category_intros/configuration-files.md @@ -0,0 +1,21 @@ +Settings that change between deployments belong in the environment, whatever Python configuration library you use. pydantic-settings reads them as typed fields. + +How to choose: + +- Typed, validated settings from environment variables and `.env` files: pydantic-settings +- An INI file your users edit, with nothing to install: configparser +- A `.env` file loaded into the environment during development: python-dotenv +- Research code that composes its config and overrides it from the command line: Hydra +- Layered settings per environment, or settings for Django and Flask: Dynaconf + +pydantic-settings loads [a settings class from environment variables or secrets files](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/). Subclass `BaseSettings` and declare each setting as a type-hinted field. [Any field you don't pass to the initializer gets its value from the environment](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/#usage), or its default when the variable isn't set. To read a `.env` file too, [set `env_file` in `model_config`](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/#dotenv-env-support), and its values get validated like the rest. + +configparser comes with Python. It reads [a basic configuration language structured like Windows INI files](https://docs.python.org/3/library/configparser.html), so end users can customize your program by editing a file. Read values with mapping access, `config['section']['option']`, which the docs [prefer for new projects](https://docs.python.org/3/library/configparser.html#legacy-api-examples) over the legacy get and set methods. Values always come back as strings, so convert them with [`getint()`, `getfloat()`, and `getboolean()`](https://docs.python.org/3/library/configparser.html#supported-datatypes). + +python-dotenv is for apps that [take their configuration from environment variables](https://saurabh-kumar.com/python-dotenv/#getting-started). In development, setting each one yourself isn't practical, so call `load_dotenv()` before the rest of your code. It reads a `.env` file when there is one and adds its values to `os.environ`, so your code reads them with `os.getenv()` as if they came from the real environment. + +Hydra is a framework for [research and other complex applications](https://hydra.cc/docs/1.3/intro/). It builds a hierarchical configuration by composition, and you override it from config files and the command line. Decorate your entry point with [`@hydra.main()`](https://hydra.cc/docs/1.3/intro/#basic-example) and point it at a YAML config. Then change a value with an argument like `db.user=root`. For options you switch between, like two databases, [create a config group](https://hydra.cc/docs/1.3/intro/#composition-example) with one file per option, and pick one in the config's `defaults` list. With [`--multirun`](https://hydra.cc/docs/1.3/intro/#multirun), one command runs your function once for each configuration you list. + +Dynaconf reads settings from [files in several formats, environment variables, and Vault or Redis](https://www.dynaconf.com/#features). Start with [`dynaconf init -f toml`](https://www.dynaconf.com/#using-python-only), which creates `config.py`, `settings.toml`, and `.secrets.toml`. Then import `settings` from `config` in your code. TOML is its default and most recommended format. To give development and production their own sections in one file, [set `environments=True`](https://www.dynaconf.com/#layered-environments-on-files). In a Django app, `django.conf.settings` [becomes a Dynaconf settings object](https://www.dynaconf.com/#using-django), and in Flask, `app.config` does. + +Keep settings that change between deployments out of your code. The [twelve-factor app stores config in environment variables](https://12factor.net/config), and python-dotenv and Dynaconf both cite it. python-dotenv, pydantic-settings, and Dynaconf let the real environment win over files. `load_dotenv()` [doesn't override variables already set](https://saurabh-kumar.com/python-dotenv/#getting-started), pydantic-settings [ranks environment variables above `.env` and secrets files](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/#field-value-priority), and Dynaconf [prioritizes environment variables over files](https://www.dynaconf.com/envvars/). Keep secrets out of Git, too: add `.env` to your `.gitignore`, and `dynaconf init` adds `.secrets.*` there for you. diff --git a/website/data/category_intros/cryptography.md b/website/data/category_intros/cryptography.md new file mode 100644 index 00000000..21f803b3 --- /dev/null +++ b/website/data/category_intros/cryptography.md @@ -0,0 +1,19 @@ +Don't touch a Python cryptography library's raw ciphers: encrypt with cryptography's Fernet or PyNaCl. Paramiko runs SSH, and ItsDangerous signs tokens. + +How to choose: + +- Encrypting data, X.509 certificates, and other standard formats: cryptography +- Encryption and signatures between your own apps, with algorithms picked for you: PyNaCl +- Hashing passwords for storage: PyNaCl +- SSH clients and servers, and SFTP: Paramiko +- Signed links, tokens, and cookies that users may read but not change: ItsDangerous + +cryptography has 2 layers: recipes that need few decisions, and low-level primitives it calls the "hazmat" layer. Its docs [recommend the recipes layer whenever possible](https://cryptography.io/en/stable/#layout), and hazmat only when necessary. For data you encrypt with a key, the recipe is [Fernet](https://cryptography.io/en/stable/fernet/): a message encrypted with it can't be read or changed without the key. To encrypt with a password, run it through a key derivation function first. The docs [recommend Argon2id](https://cryptography.io/en/stable/fernet/#using-passwords-with-fernet), and you keep the salt to derive the same key again. When Fernet doesn't fit, the docs point you to [authenticated encryption](https://cryptography.io/en/stable/hazmat/primitives/symmetric-encryption/) before any raw cipher, since encryption alone keeps data secret but doesn't stop tampering. The recipes also cover [X.509 certificates](https://cryptography.io/en/stable/x509/tutorial/), from signing requests to self-signed certs. + +PyNaCl is a binding to libsodium, a fork of NaCl. cryptography's FAQ explains the split: cryptography is [general purpose and interoperable with existing systems](https://cryptography.io/en/stable/faq/#how-does-cryptography-compare-to-nacl-networking-and-cryptography-library), while NaCl gives you a set of hand-selected algorithms. If you prefer NaCl's design, the FAQ recommends PyNaCl. With a shared key, encrypt through `SecretBox` or `Aead` and let PyNaCl [generate a random nonce](https://pynacl.readthedocs.io/en/latest/secret/#nacl.secret.Aead.encrypt) for each message, which its docs strongly recommend. Between 2 parties, `Box` [authenticates both sides](https://pynacl.readthedocs.io/en/latest/public/#nacl-public-box), and `SealedBox` sends a message only the recipient can decrypt, without proving who sent it. To store passwords, `nacl.pwhash.str()` [hashes with argon2id by default](https://pynacl.readthedocs.io/en/latest/password_hashing/#password-storage-and-verification), and `nacl.pwhash.verify()` checks a password against the hash. + +Paramiko implements SSHv2 in pure Python, as both client and server. For common client jobs like running remote commands or transferring files, its homepage [recommends a higher-level library built on it](https://www.paramiko.org/). Use Paramiko directly for low-level primitives or to run an SSH server in Python. Its client API [starts with `SSHClient`](https://docs.paramiko.org/en/latest/), and checking the server's host key is your job: call `load_system_host_keys()` before `connect()`. A host key it can't find is [rejected by default](https://docs.paramiko.org/en/latest/api/client.html#paramiko.client.SSHClient.set_missing_host_key_policy) with an `SSHException`. `open_sftp()` opens an SFTP session on the same connection. + +ItsDangerous signs data, so you can send it somewhere untrusted and get it back: the receiver [can see the data but can't modify it](https://itsdangerous.palletsprojects.com/en/stable/) without your key. Flask's default sessions are [cookies signed with it](https://flask.palletsprojects.com/en/stable/api/#flask.sessions.SecureCookieSessionInterface). Signing hides nothing, so when the data must stay secret, encrypt it with Fernet instead. Its docs say you'll [typically want a serializer, not a signer](https://itsdangerous.palletsprojects.com/en/stable/concepts/#serializer-vs-signer). `URLSafeSerializer` turns data into a URL-safe string, and `URLSafeTimedSerializer` rejects tokens older than the `max_age` you pass to `loads()`. Give each use its own [salt](https://itsdangerous.palletsprojects.com/en/stable/concepts/#the-salt), so an activation link's signature won't pass as an upgrade link's. + +Whichever pick you use, keep secret keys out of your code. ItsDangerous's docs say the key [shouldn't be saved in source code or committed to version control](https://itsdangerous.palletsprojects.com/en/stable/concepts/#the-secret-key), and suggest reading it from an environment variable. When you generate a key or token yourself, use `os.urandom()`, [never the `random` module](https://cryptography.io/en/stable/random-numbers/), which isn't cryptographically secure. Plan for rotation, too: cryptography's [`MultiFernet`](https://cryptography.io/en/stable/fernet/#cryptography.fernet.MultiFernet) and ItsDangerous both [take a list of keys](https://itsdangerous.palletsprojects.com/en/stable/concepts/#key-rotation), so old tokens keep working after you add a new key. diff --git a/website/data/category_intros/data-analysis.md b/website/data/category_intros/data-analysis.md new file mode 100644 index 00000000..5d640802 --- /dev/null +++ b/website/data/category_intros/data-analysis.md @@ -0,0 +1,17 @@ +Once your data nears the size of your RAM, switch Python data analysis libraries from pandas to Polars. Once it lives in a database, Ibis runs your code there. + +How to choose: + +- Data that fits in memory: pandas +- Handing DataFrames to libraries that expect pandas: pandas +- Large data on one machine, even bigger than your RAM: Polars +- Data already in a database or warehouse: Ibis +- Code you prototype locally, then run on a warehouse or cluster: Ibis + +pandas aims to be [the fundamental high-level building block](https://pandas.pydata.org/docs/getting_started/overview.html) for practical data analysis in Python. Its DataFrame fits time series and tables with mixed column types, like an SQL table or a spreadsheet. It's built on NumPy to work with the rest of the scientific Python stack. In production code, the docs [recommend `.loc` and `.iloc` over `[]`](https://pandas.pydata.org/docs/user_guide/indexing.html) to select data. Set values in [one `.loc` call](https://pandas.pydata.org/docs/user_guide/copy_on_write.html#chained-assignment), not chained indexing like `df["foo"][mask] = 100`. pandas keeps everything in memory, so a dataset that takes a sizable share of your RAM [gets unwieldy](https://pandas.pydata.org/docs/user_guide/scale.html), and the docs point you to other libraries. + +Polars is [built for multithreaded computing on a single machine](https://docs.pola.rs/user-guide/misc/comparison/#pandas), and its stricter API leads to fewer schema bugs. It has no index: rows are known by their position, so [no index state can change what a query means](https://docs.pola.rs/user-guide/migration/pandas/#polars-does-not-have-a-multi-indexindex). Make the lazy API [your default](https://docs.pola.rs/user-guide/migration/pandas/#be-lazy): start from a function like `scan_csv()` or call `.lazy()`, then `.collect()` the result. That lets the optimizer [filter rows and pick columns while reading the data](https://docs.pola.rs/user-guide/concepts/lazy-api/). Stay eager for exploratory work, when [you don't know yet what your query will look like](https://docs.pola.rs/user-guide/concepts/lazy-api/#when-to-use-which). Write expressions inside `select`, `with_columns`, `filter`, and `group_by`, since Polars code that [looks like pandas code](https://docs.pola.rs/user-guide/migration/pandas/#key-syntax-differences) likely runs slower than it should. For data that doesn't fit in memory, run the query on the [streaming engine](https://docs.pola.rs/user-guide/concepts/streaming/), which works through it in batches. + +Ibis [compiles one Python dataframe API](https://ibis-project.org/why#how-does-ibis-work) into each backend's native language, mostly SQL, so the database or engine does the work. pandas' ecosystem page lists it for [bridging local Python and remote databases](https://pandas.pydata.org/community/ecosystem.html#ibis). Start on the default DuckDB backend, as the tutorial [recommends](https://ibis-project.org/tutorials/basics#install-ibis), then [change the connection string](https://ibis-project.org/why#scaling-up-and-out) to run the same code on PySpark, BigQuery, or Trino. Expressions are lazy: nothing runs until you call a method like `to_pandas()`, and only then does Ibis [send the compiled query](https://ibis-project.org/tutorials/coming-from/pandas) to the backend. + +You can mix all three. Polars converts a DataFrame [with `to_pandas()`](https://docs.pola.rs/api/python/stable/reference/dataframe/api/polars.DataFrame.to_pandas.html), Ibis returns results [as pandas or Polars DataFrames](https://ibis-project.org/reference/expression-tables#ibis.expr.types.relations.Table.to_pandas), and pandas [can use PyArrow](https://pandas.pydata.org/docs/user_guide/pyarrow.html) to trade data with Arrow-based libraries like Polars. Do the heavy work in Polars or Ibis, and hand a pandas DataFrame to the libraries that expect one. diff --git a/website/data/category_intros/data-ingestion-etl.md b/website/data/category_intros/data-ingestion-etl.md new file mode 100644 index 00000000..5e919e34 --- /dev/null +++ b/website/data/category_intros/data-ingestion-etl.md @@ -0,0 +1,27 @@ +Loading data from APIs and databases into a warehouse with a Python ETL library is dlt's job. Fetching stock prices for your own research is yfinance's. + +How to choose: + +- Data from APIs and databases into a warehouse or data lake: dlt +- Stock prices and financials for your own research: yfinance +- pandas DataFrames in and out of AWS services: AWS SDK for pandas (awswrangler) +- One pipeline for both batch data and live streams: Pathway +- Chinese market data: AKShare +- SEC filings and XBRL financial statements: EdgarTools +- Many data providers behind one API: OpenBB + +dlt [loads data from messy sources into well-structured datasets](https://dlthub.com/docs/intro), and infers the schema and data types for you. For an API, start with `dlt init rest_api duckdb`: you [declare the endpoints, pagination, and authentication](https://dlthub.com/docs/dlt-ecosystem/verified-sources/rest_api/basic), and the REST API source does the rest. Build and test on DuckDB locally, then [switch the destination](https://dlthub.com/docs/reference/explainers/how-dlt-works) when you deploy. To [pick a write disposition](https://dlthub.com/docs/general-usage/incremental-loading#how-to-choose-the-right-write-disposition), ask whether your data can change: append records that never change, and merge the ones that do. Keep credentials in `secrets.toml` or environment variables, and [never commit `secrets.toml`](https://dlthub.com/docs/general-usage/credentials/setup#secretstoml-and-configtoml). + +yfinance [fetches market data from Yahoo Finance](https://github.com/ranaroussi/yfinance). Use `Ticker` for one symbol, from its price history to its financial statements, and [`download()` for several symbols](https://ranaroussi.github.io/yfinance/) at once. + +awswrangler is the AWS SDK for pandas. It [connects DataFrames to AWS data and analytics services](https://aws-sdk-pandas.readthedocs.io/en/stable/about.html) like Athena, Glue, Redshift, and S3. It [leaves credentials to boto3 sessions](https://aws-sdk-pandas.readthedocs.io/en/stable/tutorials/002%20-%20Sessions.html): pass your own as `boto3_session`, or it uses the default session. Write S3 data with `wr.s3.to_parquet(..., dataset=True, database=..., table=...)`, and the dataset [goes into the Glue Catalog](https://aws-sdk-pandas.readthedocs.io/en/stable/stubs/awswrangler.s3.to_parquet.html), where Athena can query it. + +Pathway is a [Python ETL framework for stream processing](https://github.com/pathwaycom/pathway), and a Rust engine runs your Python code. To switch between batch and streaming, you [change only the data sources](https://pathway.com/developers/user-guide/introduction/batch-processing/), and the rest of the pipeline stays the same. Streaming is the standard way to run it. Data flows in once you call `pw.run()` and nothing after that call runs, so [read the results through output connectors](https://pathway.com/developers/user-guide/introduction/streaming-and-static-modes/). Pathway is under the [Business Source License](https://pathway.com/developers/user-guide/introduction/licensing-guide/), not an open-source one: production use is free within its limits, and the code converts to Apache after 4 years. + +AKShare is a [Python library for financial data](https://akshare.akfamily.xyz/introduction.html) on stocks, futures, options, funds, bonds, and more, collected from public websites as you call it. Its [stock data](https://akshare.akfamily.xyz/data/stock/stock.html) centers on China, from A-shares and B-shares to the STAR Market, plus Hong Kong and US stocks. Each dataset is one function, like `ak.stock_zh_a_hist()` for A-share daily prices. Its interfaces break when those websites change, so [upgrade AKShare before you use it](https://akshare.akfamily.xyz/installation.html). + +EdgarTools [makes SEC filings easy to access and analyze](https://edgartools.readthedocs.io/en/latest/). Start from a `Company` or a `Filing`, and `.obj()` gives you [a typed object for that form](https://github.com/dgunning/edgartools#how-it-works), with its data as pandas DataFrames. EDGAR requires an email with every request, so set your identity first, with `set_identity()` or the [`EDGAR_IDENTITY` variable](https://edgartools.readthedocs.io/en/latest/configuration/). For bulk work, [turn on local storage](https://edgartools.readthedocs.io/en/latest/guides/local-storage/), which cuts down requests and respects the SEC's rate limits. + +OpenBB's Open Data Platform [integrates proprietary, licensed, and public data sources](https://github.com/OpenBB-finance/OpenBB) and serves them to Python, a REST API, Excel, and MCP servers. Install it in a [new environment](https://docs.openbb.co/odp/python/installation), not your system Python. Pass `provider` to each query: without it, OpenBB [picks the first available provider in alphabetical order](https://docs.openbb.co/odp/python/quickstart). Most providers [need your own API key](https://docs.openbb.co/odp/python/settings/user_settings/api_keys). + +Read the data's terms before you build a product on it, and check its numbers before you trade on them. yfinance's docs say [Yahoo's API is for personal use only](https://ranaroussi.github.io/yfinance/), AKShare says its data is [only for academic research](https://github.com/akfamily/akshare#statement), and OpenBB says its data is [not necessarily accurate](https://github.com/OpenBB-finance/OpenBB#4-disclaimer). diff --git a/website/data/category_intros/data-validation.md b/website/data/category_intros/data-validation.md new file mode 100644 index 00000000..f5bd96ea --- /dev/null +++ b/website/data/category_intros/data-validation.md @@ -0,0 +1,15 @@ +Validate API input and config with Pydantic, the Python data validation library built on type hints. Use Pandera for dataframes, jsonschema for JSON Schema. + +How to choose: + +- API input, forms, and config: Pydantic +- Data checked against a JSON Schema document: jsonschema +- pandas, polars, or PySpark dataframes: Pandera + +Pydantic builds the schema from your [type hints](https://pydantic.dev/docs/validation/latest/get-started/why/) and guarantees the types of the [output, not the input](https://pydantic.dev/docs/validation/latest/concepts/models/): by default, a numeric string passed to an int field comes out as an int. Where a wrong type should raise an error instead, turn on [strict mode](https://pydantic.dev/docs/validation/latest/concepts/strict_mode/) per field or per model. [Validate incoming JSON directly](https://pydantic.dev/docs/validation/latest/concepts/performance/) instead of parsing it into a dict first. To load config from environment variables, use [Pydantic Settings](https://pydantic.dev/docs/validation/latest/concepts/pydantic_settings/). + +jsonschema is an [implementation of the JSON Schema specification](https://python-jsonschema.readthedocs.io/en/stable/), so one schema can [work across different systems and platforms](https://json-schema.org/overview/what-is-jsonschema). When you validate many instances against one schema, [create a validator for your schema's draft once and call its `validate` method](https://python-jsonschema.readthedocs.io/en/stable/validate/). Use [`iter_errors()`](https://python-jsonschema.readthedocs.io/en/stable/errors/) to report every error, not only the first. To enforce `format` keywords such as dates or emails, [hook a format checker](https://python-jsonschema.readthedocs.io/en/stable/validate/#validating-formats) into the validator. + +Pandera validates [dataframe-like objects](https://pandera.readthedocs.io/en/stable/): define a schema once and use it on pandas, polars, PySpark, and other dataframe libraries. Write the schema as a DataFrameModel class, [much like a Pydantic model](https://pandera.readthedocs.io/en/stable/dataframe_models.html), and add the `check_types()` decorator to validate at run time. To check an existing pipeline, put [`check_input()` and `check_output()`](https://pandera.readthedocs.io/en/stable/decorators.html) on its functions. To see every failure in one run instead of only the first, validate with [`lazy=True`](https://pandera.readthedocs.io/en/stable/lazy_validation.html). + +Pick by the shape of your data. Running a dataframe through a Pydantic model row by row [might not scale](https://pandera.readthedocs.io/en/stable/pydantic_integration.html) to larger datasets, so use Pandera there; a DataFrameModel can still be a field in a Pydantic model. Pydantic can [generate a JSON Schema](https://pydantic.dev/docs/validation/latest/concepts/json_schema/) from any model for tools that read the format. jsonschema works the other way: it validates data against a JSON Schema document you already have. diff --git a/website/data/category_intros/data-visualization.md b/website/data/category_intros/data-visualization.md new file mode 100644 index 00000000..3cd21ae7 --- /dev/null +++ b/website/data/category_intros/data-visualization.md @@ -0,0 +1,36 @@ +Pick your Python data visualization library by where the chart goes. Papers take Matplotlib figures, web pages take plotly charts, and data apps use Streamlit. + +How to choose: + +- Figures for papers and reports: Matplotlib +- Interactive charts on a web page: plotly +- Data dashboards and apps: Streamlit +- Statistical charts from a dataframe: seaborn +- Charts declared from dataframe columns: Vega-Altair +- Interactive plots that call back into Python: Bokeh +- Maps in any projection: Cartopy +- Graph diagrams laid out by Graphviz: PyGraphviz +- Knowledge graph of a codebase: graphify +- Demos for machine learning models: Gradio + +Matplotlib has two interfaces. For complicated plots and code you reuse, its docs [suggest the explicit, object-oriented one](https://matplotlib.org/stable/users/explain/quick_start.html#the-explicit-and-the-implicit-interfaces): create the figure with `fig, ax = plt.subplots()`, then call methods on `ax`. The implicit pyplot style is fine for quick interactive work. When you make the same plot for many datasets, write a function that takes the `ax` to draw on. + +plotly draws interactive charts in the browser with plotly.js. Start with [Plotly Express](https://plotly.com/python/plotly-express/), which its docs call the recommended starting point for most common figures: pass a DataFrame and column names, and one call builds the figure. Drop to `go.Figure` for figures Plotly Express can't make or makes awkward, like [subplots of different types or dual-axis plots](https://plotly.com/python/graph-objects/#When-to-use-Graph-Objects-vs-Plotly-Express). To share a chart, [`write_html`](https://plotly.com/python/interactive-html-export/) saves it as an HTML file that stays interactive in any browser. + +seaborn builds statistical graphics on Matplotlib and works on whole datasets. Its docs [recommend the figure-level functions](https://seaborn.pydata.org/tutorial/function_overview.html#relative-merits-of-figure-level-functions), like `relplot()`, for most plots. For one figure that combines different kinds of plots, set it up in Matplotlib and fill it in with axes-level functions. Keep your data in [long form](https://seaborn.pydata.org/tutorial/data_structure.html), one column per variable and one row per observation, as most of seaborn's examples do. + +Vega-Altair is declarative: you [link data columns to encoding channels](https://altair-viz.github.io/getting_started/overview.html) like the x-axis, y-axis, and color, and it handles the rest on top of Vega-Lite. Its docs say the API is [more limited than Matplotlib's or Bokeh's](https://altair-viz.github.io/getting_started/project_philosophy.html), a trade they make to keep exploring data simple. A chart carries its data inside the spec, so for a large dataset, [enable the VegaFusion data transformer](https://altair-viz.github.io/user_guide/large_datasets.html#vegafusion-data-transformer) or pass the data by URL. + +Bokeh builds interactive plots for the browser without any JavaScript from you. Start with [`bokeh.plotting`, its primary interface](https://docs.bokeh.org/en/latest/docs/user_guide/intro.html#the-bokeh-plotting-interface), and `output_file()` for a standalone HTML file or `output_notebook()` for Jupyter. What sets it apart is the [Bokeh server](https://docs.bokeh.org/en/latest/docs/user_guide/server/server_introduction.html), which keeps data in sync between Python and the browser. Widgets can then run Python callbacks, and plots can stream data. Write the app as a script and [serve it with `bokeh serve`](https://docs.bokeh.org/en/latest/docs/user_guide/server/app.html#building-applications). + +Cartopy draws maps on Matplotlib and [suits data over large areas](https://cartopy.readthedocs.io/stable/), where Cartesian math breaks down at the poles and the dateline. Set the map's projection on the axes, and [always pass `transform`](https://cartopy.readthedocs.io/stable/tutorials/understanding_transform.html) to say which coordinate system your data is in. + +PyGraphviz is a Python interface to Graphviz. Build a graph with `AGraph` or read a DOT file into one, then [lay it out and draw it](https://pygraphviz.github.io/documentation/stable/tutorial.html#layout-and-drawing) with one of Graphviz's layout programs. + +graphify maps a project's code, docs, PDFs, and images into a [knowledge graph your coding agent can query](https://github.com/Graphify-Labs/graphify). It draws the graph as a `graph.html` you can click through in a browser. It parses code locally, so code never leaves your machine; docs and media go through your agent's model. Install the `graphifyy` package, run `graphify install`, then type `/graphify .` in your agent. + +Streamlit turns a Python script into a data app, and [reruns the whole script from top to bottom](https://docs.streamlit.io/get-started/fundamentals/main-concepts) every time something on screen changes. Start it with `streamlit run`. To skip repeated work on those reruns, [cache](https://docs.streamlit.io/get-started/fundamentals/advanced-concepts) data with `st.cache_data`, and shared resources like ML models or database connections with `st.cache_resource`. Keep per-user values in Session State. + +Gradio wraps a Python function, often a machine learning model, in a web UI. [Use `gr.Interface`](https://gradio.app/guides/quickstart) for a demo with inputs and outputs, `gr.ChatInterface` for a chatbot, and `gr.Blocks` for custom layouts and data flows. `launch(share=True)` gives you a public link, and [Hugging Face Spaces](https://gradio.app/guides/sharing-your-app#hosting-on-hf-spaces) hosts the app for good. Anyone with the link can call your function, so [put a login in front of it](https://gradio.app/guides/sharing-your-app#password-protected-app) or keep sensitive data out. Its security docs also recommend that you [set `max_file_size` and keep `allowed_paths` as small as possible](https://gradio.app/guides/file-access#best-practices). + +seaborn and Cartopy draw on Matplotlib Axes, so to [customize what they draw](https://matplotlib.org/stable/users/explain/figure/api_interfaces.html), use Matplotlib's explicit Axes interface. Streamlit shows [Matplotlib, plotly, and Vega-Altair figures](https://docs.streamlit.io/develop/api-reference/charts), and Gradio's [`gr.Plot`](https://gradio.app/docs/gradio/plot) takes those plus Bokeh, so your plotting code carries over into an app. diff --git a/website/data/category_intros/database-drivers.md b/website/data/category_intros/database-drivers.md new file mode 100644 index 00000000..90edd10c --- /dev/null +++ b/website/data/category_intros/database-drivers.md @@ -0,0 +1,32 @@ +PostgreSQL gets Psycopg, MySQL mysqlclient, and SQLite the built-in sqlite3. Elsewhere, use the Python database driver from your database's maker. + +How to choose: + +- PostgreSQL, from sync or async code: Psycopg +- MySQL or MariaDB: mysqlclient, or PyMySQL for pure Python under the MIT license +- SQLite in your app: sqlite3 +- PostgreSQL from asyncio code, when query speed comes first: asyncpg +- Loading JSON or CSV into SQLite and reshaping its tables, from the shell or Python: sqlite-utils +- ClickHouse: ClickHouse Connect, or clickhouse-driver for the native TCP protocol +- Any other database with an ODBC driver: pyodbc +- Oracle Database: python-oracledb +- SQL Server or Azure SQL, with no driver manager to install: mssql-python +- Redis: redis-py +- MongoDB: PyMongo, or Django MongoDB Backend in a Django project +- Apache Cassandra: cassandra-driver + +Psycopg keeps the [DB-API interface](https://www.psycopg.org/psycopg3/docs/) of the older Psycopg and adds asyncio support, so sync and async code share one driver. Open a connection [in a `with` block](https://www.psycopg.org/psycopg3/docs/basic/usage.html#connection-context): it commits when the block ends, rolls back if an exception is raised, and closes the connection either way. When several threads need connections, take them from a [`ConnectionPool`](https://www.psycopg.org/psycopg3/docs/advanced/pool.html#basic-connection-pool-usage), or an `AsyncConnectionPool` in async code. Django [recommends Psycopg](https://docs.djangoproject.com/en/stable/ref/databases/#postgresql-notes) for PostgreSQL. + +asyncpg is built for asyncio and [speaks PostgreSQL's protocol natively](https://github.com/MagicStack/asyncpg) instead of hiding it behind the DB-API, so its API is its own, down to `$1` placeholders. In a server, [use its connection pool](https://magicstack.github.io/asyncpg/current/usage.html#connection-pools): take a connection per request with `async with pool.acquire()`, and wrap writes in `async with connection.transaction()`. Outside a transaction, each statement commits right away. + +mysqlclient is a native driver that builds against the MySQL client library, and it's [Django's recommended choice](https://docs.djangoproject.com/en/stable/ref/databases/#mysql-db-api-drivers) for MySQL. PyMySQL is [pure Python](https://github.com/PyMySQL/PyMySQL), so it installs without that library. The licenses differ, too: mysqlclient is [GPL](https://github.com/PyMySQL/mysqlclient/blob/main/LICENSE), and PyMySQL is [MIT](https://github.com/PyMySQL/PyMySQL/blob/main/LICENSE). + +sqlite3 ships with Python, and its docs suggest SQLite for an app's internal storage, or for [a prototype you later port](https://docs.python.org/3/library/sqlite3.html) to a larger database like PostgreSQL. Use the connection [as a context manager](https://docs.python.org/3/library/sqlite3.html#sqlite3-connection-context-manager) to commit or roll back a transaction; it doesn't close the connection, so close it yourself. sqlite-utils is [not a full ORM](https://sqlite-utils.datasette.io/en/stable/) but a set of helpers for creating a SQLite database and filling it with data, from Python or its command line. [Pipe JSON or CSV into it](https://github.com/simonw/sqlite-utils), and it creates the table for you. It also runs schema changes that SQLite's `ALTER TABLE` can't, like changing a column's type. + +ClickHouse Connect is the Python driver that [ClickHouse's own docs](https://clickhouse.com/docs/integrations/language-clients/python/index) cover. It has a sync and an async client over the HTTP interface, which works through load balancers and proxies. clickhouse-driver talks ClickHouse's [native TCP protocol](https://clickhouse-driver.readthedocs.io/en/latest/) instead. To insert rows fast with it, [pass them separately](https://clickhouse-driver.readthedocs.io/en/latest/quickstart.html#inserting-data) and end the statement with `VALUES`. + +pyodbc connects to [any database with an ODBC driver](https://github.com/mkleehammer/pyodbc). On macOS and Linux, install an ODBC driver manager like unixODBC first; Windows has one built in. python-oracledb is Oracle's own driver and the [successor to cx_Oracle](https://python-oracledb.readthedocs.io/en/latest/user_guide/introduction.html). Its default Thin mode [connects without Oracle Client libraries](https://python-oracledb.readthedocs.io/en/latest/user_guide/initialization.html), which is enough for most apps. mssql-python is [Microsoft's own driver](https://learn.microsoft.com/en-us/sql/connect/python/mssql-python/migrate-from-pyodbc) for SQL Server and Azure SQL. It [connects without an external driver manager](https://learn.microsoft.com/en-us/sql/connect/python/mssql-python/python-sql-driver-mssql-python) and [pools connections by default](https://github.com/microsoft/mssql-python#connection-pooling). + +redis-py is [the Python client for Redis](https://redis.io/docs/latest/develop/clients/redis-py/), with asyncio support. PyMongo is [the recommended way](https://www.mongodb.com/docs/languages/python/pymongo-driver/current/) to work with MongoDB from Python: use `MongoClient` in sync code and [`AsyncMongoClient`](https://www.mongodb.com/docs/languages/python/pymongo-driver/current/connect/mongoclient/) in async code. Django MongoDB Backend is [a Django database backend](https://github.com/mongodb/django-mongodb-backend) that uses PyMongo, so your Django models live in MongoDB. Joins [don't perform well on large tables](https://django-mongodb-backend.readthedocs.io/en/latest/faq/#performance) there, so model related data as embedded models. With cassandra-driver, [use prepared statements](https://docs.datastax.com/en/developer/python-driver/latest/getting_started/#prepared-statement) for queries you run often, so Cassandra doesn't parse them again each time. + +Whatever the database, create the client or pool once per process and share it. ClickHouse Connect's docs say to [create clients once at startup](https://clickhouse.com/docs/integrations/language-clients/python/driver-api#client-lifecycle-and-best-practices), and a PyMongo `MongoClient` or a redis-py `Redis` object already holds a pool. If your server forks worker processes, [create the client or pool after the fork](https://www.psycopg.org/psycopg3/docs/advanced/async.html#concurrent-operations). Pass values as query parameters. Building SQL strings yourself [opens the door to SQL injection](https://docs.python.org/3/library/sqlite3.html#sqlite3-placeholders). diff --git a/website/data/category_intros/database.md b/website/data/category_intros/database.md new file mode 100644 index 00000000..1820d93a --- /dev/null +++ b/website/data/category_intros/database.md @@ -0,0 +1,27 @@ +No server to run: a Python database library can live inside your process. Pick DuckDB when you run analytical SQL, LanceDB when you search vectors. + +How to choose: + +- Analytical SQL over DataFrames and Parquet files: DuckDB +- Vector search with the raw data stored beside its embeddings: LanceDB +- ClickHouse SQL in a notebook, the same queries your ClickHouse server runs: chDB +- A RAG prototype that embeds your text for you: Chroma +- Vector search, full-text search, and filters in one query: Zvec +- Tables whose columns run AI models on every insert: Pixeltable +- Python dicts in a JSON file, for a small app with one process: TinyDB + +DuckDB is built for [analytical queries](https://duckdb.org/why_duckdb), and it runs inside your Python process with no database server to install. `duckdb.sql()` runs on an in-memory database, while [`duckdb.connect()` with a file name](https://duckdb.org/docs/current/clients/python/overview#persistent-storage) keeps your tables on disk. It [queries pandas and Polars DataFrames and Arrow tables directly](https://duckdb.org/docs/current/clients/python/overview#dataframes), and Parquet files by their file name. In a package others import, [create your own connection objects](https://duckdb.org/docs/current/clients/python/overview#connection-object-and-module) instead of calling the module's functions, which share one global database. + +chDB is an [in-process SQL engine powered by ClickHouse](https://clickhouse.com/docs/chdb), for ClickHouse SQL without a ClickHouse server. It speaks [the full ClickHouse SQL dialect](https://clickhouse.com/resources/engineering/what-is-chdb), so a query you write in a notebook runs unchanged on a ClickHouse server later. `chdb.query()` is stateless. For tables that last across queries, open a `Session`, and [give it a directory name](https://clickhouse.com/docs/chdb/getting-started#creating-a-table-from-json-file) to keep them on disk. + +Chroma's embedded mode is for [prototyping and experimentation](https://docs.trychroma.com/reference/architecture/overview). It gets a RAG prototype running fast, since it [embeds and indexes your documents for you](https://docs.trychroma.com/docs/overview/getting-started) with a default model that runs on your machine. Save data to a directory with [`PersistentClient`](https://docs.trychroma.com/docs/run-chroma/clients#persistent-client). For production, its docs [prefer a Chroma server](https://docs.trychroma.com/reference/python/client#persistentclient) that your app connects to as a client. + +LanceDB is an [embedded retrieval library](https://docs.lancedb.com/) that runs in your process: [point it at a local directory](https://docs.lancedb.com/quickstart#connect-via-local-directory-path), or at an object storage URI like `s3://`. It [stores the raw data, metadata, and embeddings together](https://docs.lancedb.com/faq/faq-oss#what-makes-lancedb-different), and its indexes live on disk. + +Zvec is a vector database that [runs entirely in-process](https://zvec.org/en/docs/db/), from notebooks and servers to edge devices. [Define a schema](https://zvec.org/en/docs/db/quickstart/) with scalar fields and vectors, then create a collection from it. One query can [combine vector similarity, full-text search, and filters](https://github.com/alibaba/zvec#user-content--features). + +Pixeltable is more than a vector store: it's [the database, orchestration, and serving](https://docs.pixeltable.com/overview/pixeltable) in one Python file. [Model inference can go in a computed column](https://docs.pixeltable.com/tutorials/computed-columns), which [runs on insert and on update](https://docs.pixeltable.com/overview/how-it-works). Each new row gets its model outputs without a pipeline you rerun. Put an embedding index on a column, and [each insert keeps it current](https://docs.pixeltable.com/howto/coming-from), with no separate vector database. + +TinyDB is pure Python with no dependencies, and [`TinyDB('db.json')`](https://tinydb.readthedocs.io/en/latest/getting-started.html) gives you a database that stores Python dicts in that JSON file. Its docs call it [the wrong database](https://tinydb.readthedocs.io/en/latest/intro.html#why-not-use-tinydb) when you need access from several processes or threads, indexes, ACID guarantees, or high performance. + +Most of these databases expect one process to write at a time. DuckDB lets [one process read and write](https://duckdb.org/docs/current/connect/concurrency#single-process), or several processes only read. A chDB data directory [opens in one process at a time](https://clickhouse.com/docs/chdb/getting-started). Zvec shares a collection across processes [in read-only mode](https://zvec.org/en/docs/db/collections/open/), and Pixeltable keeps [one writer process](https://docs.pixeltable.com/howto/deployment/operations) even with several API workers. diff --git a/website/data/category_intros/date-and-time.md b/website/data/category_intros/date-and-time.md new file mode 100644 index 00000000..eeda6507 --- /dev/null +++ b/website/data/category_intros/date-and-time.md @@ -0,0 +1,21 @@ +Most code should handle time zones with zoneinfo and pick python-dateutil, the Python date library for parsing date strings and adding months. + +How to choose: + +- Time zones by IANA name, like Europe/Paris: zoneinfo +- Parsing date strings, adding months, and recurring dates: python-dateutil +- Dates as people write them, like "3 days ago", in many languages: dateparser +- An easier datetime API whose objects are still datetimes: pendulum +- Exact and local times as separate types, with DST-safe math: whenever + +zoneinfo brings the IANA time zone database to Python, and the datetime docs say [its usage is recommended](https://docs.python.org/3/library/datetime.html#tzinfo-objects). Attach a ZoneInfo to a datetime [through the constructor, `replace()`, or `astimezone()`](https://docs.python.org/3/library/zoneinfo.html#using-zoneinfo). Some systems, Windows among them, have no IANA database, so if your code runs across platforms, [declare a dependency on tzdata](https://docs.python.org/3/library/zoneinfo.html#data-sources). + +python-dateutil adds [extensions to the standard datetime module](https://dateutil.readthedocs.io/en/stable/); install it as `python-dateutil` and import it as `dateutil`. Its parser [reads most known formats](https://dateutil.readthedocs.io/en/stable/parser.html) and returns a datetime even for an ambiguous date. For input like 01/05/09, set `dayfirst` or `yearfirst` to match your data. For calendar math, `relativedelta` takes [plural arguments that add and singular ones that replace](https://dateutil.readthedocs.io/en/stable/relativedelta.html): `months=+1` moves a month ahead, and `day=1` jumps to the first. `rrule` builds [recurring dates from iCalendar rules](https://dateutil.readthedocs.io/en/stable/rrule.html). + +dateparser reads relative dates like "two weeks ago" and absolute ones in more than 200 language locales. Its docs say it [stands out](https://dateparser.readthedocs.io/en/latest/#common-use-cases) for scraped pages, logs, and other data from mixed sources, and for letting users type dates in their own words. Call `dateparser.parse()`, and [pass `languages` when you know them](https://dateparser.readthedocs.io/en/latest/#how-to-use), so it skips language detection. When you parse many dates from one source, [use `DateDataParser`](https://dateparser.readthedocs.io/en/latest/usage.html), which remembers the languages it has found. + +pendulum's classes are [drop-in replacements for the native ones](https://pendulum.eustace.io/docs/#introduction), since they inherit from datetime. Every instance is time zone aware and in UTC by default. Its docs call aware datetimes [the preferred and recommended way](https://pendulum.eustace.io/docs/#instantiation) to use it. For tests, install `pendulum[test]` and [travel in time](https://pendulum.eustace.io/docs/#testing). + +whenever puts exact time and local time in [separate types](https://whenever.readthedocs.io/en/latest/guide/choosing-a-type.html): an instant when only the moment matters, a zoned datetime when the local time matters too. Mixing up naive and aware [becomes a type error](https://whenever.readthedocs.io/en/latest/), and DST is handled in all arithmetic. A standard datetime [does no time zone adjustment](https://docs.python.org/3/library/datetime.html#datetime-objects) when you add a timedelta to it. In production, [turn whenever's DST warnings into errors](https://whenever.readthedocs.io/en/latest/faq.html#why-warnings-instead-of-errors) with Python's standard warnings filter. + +Decide whether you extend datetime or replace it. zoneinfo, python-dateutil, and dateparser all use standard datetime objects, so they work together. pendulum's objects are datetimes too, but code that checks the exact type, like sqlite3 and some database drivers, [needs an adapter registered](https://pendulum.eustace.io/docs/#limitations). whenever [doesn't subclass datetime at all](https://whenever.readthedocs.io/en/latest/faq.html#why-no-drop-in-replacement-for-datetime), so [convert to and from standard datetimes](https://whenever.readthedocs.io/en/latest/guide/stdlib-convert.html) where other code needs one. diff --git a/website/data/category_intros/debugging-tools.md b/website/data/category_intros/debugging-tools.md new file mode 100644 index 00000000..f40a28cc --- /dev/null +++ b/website/data/category_intros/debugging-tools.md @@ -0,0 +1,33 @@ +Before you add another print call, try a Python debugging tool. Step through your code in ipdb, and when it's slow, find where the time goes with py-spy. + +How to choose: + +- Stepping through code with pdb's commands: ipdb +- Profiling without changing code, even in production: py-spy +- A full-screen debugger in the terminal: PuDB +- Tracing calls through a big application: Hunter +- Memory use and leaks, on Linux and macOS: Memray +- A call tree that includes time spent waiting on I/O: pyinstrument +- CPU, GPU, and memory profiles line by line: Scalene +- Debug panels in a Django or Flask app: Django Debug Toolbar or Flask-DebugToolbar +- Print debugging that shows each expression: IceCream + +ipdb gives you the IPython debugger with [the same interface as pdb](https://github.com/gotcha/ipdb). It adds tab completion, syntax highlighting, and better tracebacks. Call `ipdb.set_trace()` where you want to stop and look around. + +PuDB is a full-screen debugger that runs in your terminal, with the source, the stack, breakpoints, and variables [all visible at once](https://documen.tician.de/pudb/). Call `from pudb import set_trace; set_trace()` where you want to stop, or [run a whole script](https://documen.tician.de/pudb/starting.html) under it with `python -m pudb my-script.py`. + +Hunter traces what your code does, to help you [understand and debug big applications](https://github.com/ionelmc/python-hunter). Its main selling point is filtering the events you see. Start it from code with `hunter.trace()`, from the `PYTHONHUNTER` environment variable, or with the `hunter-trace` CLI, which [attaches to a running process](https://python-hunter.readthedocs.io/en/latest/introduction.html#activation). To see only your own code, [set `stdlib=False`](https://python-hunter.readthedocs.io/en/latest/cookbook.html#typical). + +py-spy shows where your program spends its time [without restarting it or changing its code](https://github.com/benfred/py-spy). It runs outside your program's process, so its docs call it safe to use on production code. `py-spy record -o profile.svg --pid 12345` writes a flame graph of a running process, or pass `-- python myprogram.py` in place of the PID to start one. When a program hangs, `py-spy dump` prints its current call stack. + +pyinstrument records [wall-clock time](https://pyinstrument.readthedocs.io/en/latest/how-it-works.html#wall-clock-time-not-cpu-time), so the time your program spends downloading data, reading files, and talking to databases shows up in its call tree. It samples the call stack instead of tracing every call, which [keeps its overhead low](https://pyinstrument.readthedocs.io/en/latest/how-it-works.html#statistical-profiling-not-tracing). Type `pyinstrument script.py` instead of `python script.py`, or [wrap the code you want to profile](https://pyinstrument.readthedocs.io/en/latest/guide.html#profile-a-specific-chunk-of-code) in a `with pyinstrument.profile():` block. + +Scalene profiles CPU, GPU, and memory [line by line](https://github.com/plasma-umass/scalene). It separates the time spent in Python from the time spent in native code, so you can focus on the code you can actually improve. It also points to the lines responsible for memory growth and likely leaks. + +Memray tracks memory allocations [in Python code, native extension modules, and the interpreter itself](https://bloomberg.github.io/memray/overview.html). It traces every function call rather than sampling, so the call stacks it reports are accurate. Use it to find what's using memory, where it leaks, and which code allocates the most. Profile [in two steps](https://bloomberg.github.io/memray/getting_started.html): `memray run example.py` saves the allocations to a file, and `memray flamegraph` turns that file into a report. Memray only works on Linux and macOS, so on Windows, profile memory with Scalene. + +Django Debug Toolbar adds panels with debug information about the current request and response. [Set it up](https://django-debug-toolbar.readthedocs.io/en/latest/installation.html) by adding its app, URLs, and middleware. The toolbar shows only for the IP addresses in `INTERNAL_IPS`, so add `"127.0.0.1"` there. Its docs warn that it [isn't hardened for production](https://django-debug-toolbar.readthedocs.io/en/latest/configuration.html#show-toolbar-callback) or public servers. Flask-DebugToolbar is [a port of it](https://github.com/pallets-eco/flask-debugtoolbar) for Flask: pass your app to `DebugToolbarExtension(app)`, and the toolbar [appears in HTML responses when debug mode is on](https://flask-debugtoolbar.readthedocs.io/en/latest/#usage). + +IceCream's `ic()` is [like `print()`, but better](https://github.com/gruns/icecream): `ic(foo(123))` prints both the expression and its value: `ic| foo(123): 456`. With no arguments, it prints the file, line number, and function it's called from. It returns its arguments, so you can wrap it around code that's already there. When you're done, `ic.disable()` turns off all its output. + +ipdb and PuDB both keep the interface of pdb, the debugger that ships with Python. You don't have to import either one in your code. Call the built-in `breakpoint()` instead, and set the `PYTHONBREAKPOINT` environment variable to [the function it should run](https://docs.python.org/3/using/cmdline.html#envvar-PYTHONBREAKPOINT), like `ipdb.set_trace` or `pudb.set_trace`. Left unset, `breakpoint()` starts pdb. diff --git a/website/data/category_intros/deep-learning.md b/website/data/category_intros/deep-learning.md new file mode 100644 index 00000000..9201d434 --- /dev/null +++ b/website/data/category_intros/deep-learning.md @@ -0,0 +1,24 @@ +Three Python deep learning frameworks, three strengths: PyTorch for research and new architectures, Keras for a high-level API, JAX for compiled code on TPUs. + +How to choose: + +- Research and new model architectures: PyTorch +- High-level model building on JAX or PyTorch: Keras +- NumPy-style array code compiled for GPUs and TPUs: JAX +- PyTorch training without the loop boilerplate: PyTorch Lightning +- Reinforcement learning environments: Gymnasium +- Reinforcement learning algorithms: Stable-Baselines3 + +PyTorch models are `nn.Module` subclasses, and autograd [builds the computational graph as your code runs](https://docs.pytorch.org/docs/stable/user_guide/pytorch_main_components.html). Install it with [the command from its selector](https://pytorch.org/get-started/locally/), which matches your OS and GPU. To keep a model, [save its `state_dict`](https://docs.pytorch.org/tutorials/beginner/saving_loading_models.html#save-load-state-dict-recommended) instead of pickling the whole module. Wrap your model in [`torch.compile`](https://docs.pytorch.org/tutorials/intermediate/torch_compile_tutorial.html) to speed it up with minimal code changes. To train on more than one GPU, use [DistributedDataParallel](https://docs.pytorch.org/docs/stable/notes/cuda.html#use-nn-parallel-distributeddataparallel-instead-of-multiprocessing-or-nn-dataparallel). + +Keras is a [multi-framework API](https://keras.io/getting_started/about/): a Keras model can run as a PyTorch Module or as a JAX function. Install a backend next to it and [set `KERAS_BACKEND`](https://keras.io/getting_started/#configuring-your-backend) before you import Keras. Build simple models as a Sequential stack of layers and anything more complex with the functional API. Save the whole model to [a `.keras` file](https://keras.io/getting_started/faq/#what-are-my-options-for-saving-models) with `model.save()` instead of pickling it. The file [reloads with any backend](https://keras.io/keras_3/). + +JAX does [accelerator-oriented array computation](https://docs.jax.dev/en/latest/) with a NumPy-style API and composable transformations: `jax.grad` for derivatives, `jax.jit` for compilation, and `jax.vmap` for batching. The transformations only work on [functionally pure functions](https://docs.jax.dev/en/latest/notebooks/Common_Gotchas_in_JAX.html#pure-functions), so pass all data in as arguments and return every result. JAX itself stays narrow: to train neural networks, use the [JAX AI Stack](https://docs.jaxstack.ai/en/latest/getting_started.html), with Flax NNX for models and Optax for optimizers. On NVIDIA GPUs, the JAX team strongly recommends [installing CUDA and cuDNN from pip wheels](https://docs.jax.dev/en/latest/installation.html#pip-installation-nvidia-gpu-cuda-installed-via-pip-easier). + +PyTorch Lightning [organizes PyTorch code to remove boilerplate](https://lightning.ai/docs/pytorch/stable/home/introduction): you write the model logic in a LightningModule, and the Trainer handles devices, precision, and distributed training. Install it as [the `lightning` package](https://lightning.ai/docs/pytorch/stable/home/installation). Its [style guide](https://lightning.ai/docs/pytorch/stable/reference/starter/style_guide) recommends keeping each LightningModule self-contained, the model separate from the system that trains it, and data loading in a LightningDataModule. To keep your own training loop, use [Lightning Fabric](https://lightning.ai/docs/fabric/stable), which scales a plain PyTorch script after you change a few lines. + +Gymnasium is [an API standard for reinforcement learning](https://gymnasium.farama.org/), with a collection of reference environments. It's the maintained fork of OpenAI's Gym, and many older tutorials still use Gym's old API, so follow its [migration guide](https://gymnasium.farama.org/introduction/migration_guide/) when you port one. [Register your own environment](https://gymnasium.farama.org/introduction/create_custom_env/#registering-and-making-the-environment) so `gymnasium.make()` creates it like a built-in one, and run [`check_env`](https://gymnasium.farama.org/introduction/create_custom_env/#check-environment-validity) on it to catch common issues. + +Stable-Baselines3 is a set of [reliable implementations of reinforcement learning algorithms in PyTorch](https://stable-baselines3.readthedocs.io/en/master/), and it trains on any environment that [follows the Gymnasium interface](https://stable-baselines3.readthedocs.io/en/master/guide/custom_env.html). It [assumes you know some reinforcement learning](https://github.com/DLR-RM/stable-baselines3). Its tips page recommends [starting from the RL Zoo's tuned hyperparameters](https://stable-baselines3.readthedocs.io/en/master/guide/rl_tips.html#general-advice-when-using-reinforcement-learning) and normalizing the agent's input. Evaluate the agent on [a separate test environment](https://stable-baselines3.readthedocs.io/en/master/guide/rl_tips.html#how-to-evaluate-an-rl-algorithm), since training adds exploration noise. [Pick an algorithm](https://stable-baselines3.readthedocs.io/en/master/guide/rl_tips.html#which-algorithm-should-i-use) by your action space first: DQN handles only discrete actions, and SAC only continuous ones. + +PyTorch's security policy says [running untrusted models is equivalent to running untrusted code](https://github.com/pytorch/pytorch/blob/main/SECURITY.md), so run untrusted ones in a sandbox. Load checkpoints with [`weights_only=True`](https://docs.pytorch.org/docs/stable/notes/serialization.html#weights-only-security) in `torch.load`, and leave [`safe_mode`](https://keras.io/api/models/model_saving_apis/model_saving_and_loading/) on when Keras loads a model. diff --git a/website/data/category_intros/devops-tools.md b/website/data/category_intros/devops-tools.md new file mode 100644 index 00000000..72efeb34 --- /dev/null +++ b/website/data/category_intros/devops-tools.md @@ -0,0 +1,46 @@ +Across a fleet of servers, Ansible applies YAML playbooks over SSH with no agent to install. Clouds publish their own Python DevOps tools, like Boto3 for AWS. + +How to choose: + +- Servers configured from YAML playbooks over SSH: Ansible +- AWS, Azure, or Google Cloud from Python code: Boto3, the Azure SDK for Python, or google-cloud-python +- AWS from your terminal and shell scripts: AWS CLI +- A new cloud instance set up on its first boot: cloud-init +- Server configuration written in Python instead of YAML: pyinfra +- An agent on every server, run from a central master: Salt +- Shell commands on remote servers, run from Python code: Fabric +- Python APIs and event handlers on AWS Lambda: Chalice +- Process and system stats, or other programs called as functions: psutil, sh +- Your app's errors, or Celery workers and tasks, monitored live: Sentry SDK, Flower +- Your app's processes kept running on Unix: Supervisor +- Encrypted backups, or chaos engineering experiments: BorgBackup, Chaos Toolkit + +Ansible playbooks [declare the state you want each system in](https://docs.ansible.com/projects/ansible/latest/getting_started/introduction.html), written in YAML. Ansible connects over SSH with your existing credentials, so the servers it manages need no extra software. When a system already matches the playbook, Ansible changes nothing. Run a playbook with `ansible-playbook`, and run it [with `--check` first](https://docs.ansible.com/projects/ansible/latest/playbook_guide/playbooks_intro.html#running-playbooks-in-check-mode) to get a report of the changes it would make, without making them. + +Boto3 is the AWS SDK for Python, and it [shares its low-level core with the AWS CLI](https://docs.aws.amazon.com/boto3/latest/guide/quickstart.html). The AWS CLI calls the same AWS APIs from your shell, for exploring a service and writing shell scripts. [Install it from AWS's own installers](https://docs.aws.amazon.com/cli/latest/userguide/cli-chap-welcome.html): the builds in package managers are unofficial. + +The Azure SDK for Python is a set of separate libraries for specific Azure services. Its [management libraries, named `azure-mgmt-*`](https://learn.microsoft.com/en-us/azure/developer/python/sdk/azure-sdk-overview#create-and-manage-azure-resources-with-management-libraries), create and configure resources, while its client libraries work with resources that already exist. On Google Cloud, google-cloud-python holds the [Cloud Client Libraries, the option Google recommends](https://docs.cloud.google.com/apis/docs/client-libraries-explained) for calling its APIs from code. + +cloud-init gives a new cloud instance [its configuration on first boot](https://docs.cloud-init.io/en/latest/), with nothing to install, and every major public cloud supports it. Write that configuration as [cloud-config](https://docs.cloud-init.io/en/latest/explanation/format/cloud-config.html), YAML whose keys describe the state you want, like packages, users, and SSH keys. For more complex configuration, cloud-init [can hand over to a tool like Ansible](https://docs.cloud-init.io/en/latest/explanation/introduction.html). + +pyinfra [turns Python code into shell commands and runs them on your servers](https://github.com/pyinfra-dev/pyinfra): think Ansible, but Python instead of YAML. A deploy is [an `inventory.py` of hosts and a `deploy.py` of operations](https://docs.pyinfra.com/en/latest/getting-started.html), run with `pyinfra inventory.py deploy.py`. Operations declare a state, like a package being installed, and pyinfra changes only what differs. The target hosts need nothing but an SSH server. + +Salt is [a remote execution framework for configuration management and orchestration](https://docs.saltproject.io/salt/user-guide/en/latest/topics/overview.html). A Salt master sends commands to minions: the systems it manages, each running the salt-minion service. salt-ssh reaches systems without that agent. Still, Salt's docs [recommend the standard install](https://docs.saltproject.io/salt/install-guide/en/latest/topics/overview.html#standard-installation-overview) of a master plus minions for most organizations, since the agentless setup lacks some features. + +Fabric is [a library that runs shell commands over SSH](https://www.fabfile.org/) and returns the results as Python objects, built on Invoke and Paramiko. Open a `Connection` to a host and call `run()` on it. To run your code from the shell, [write `@task` functions in a `fabfile.py`](https://docs.fabfile.org/en/latest/getting-started.html#addendum-the-fab-command-line-tool) and call them with `fab`. + +Chalice is [a framework for serverless apps on AWS](https://aws.github.io/chalice/). Flask-style decorators hook your functions up to HTTP routes, schedules, and S3 events. Then [`chalice deploy`](https://aws.github.io/chalice/quickstart.html) provisions what they need on API Gateway and Lambda. + +psutil [reads process and system stats](https://psutil.io/), like CPU, memory, disks, and network, with one API on every platform it supports. Its docs call [parsing the output of `ps` or `top`](https://psutil.io/alternatives/) fragile, since psutil reads the same kernel data directly. sh [calls any program as if it were a function](https://sh.readthedocs.io/en/latest/), on Unix-like systems only. + +The Sentry SDK reports your app's errors and uncaught exceptions to Sentry. [Initialize it in your app's entry point](https://docs.sentry.io/platforms/python/), as early as possible. For Django, FastAPI, or another web framework, follow that framework's guide instead. + +Supervisor [monitors and controls your project's processes](https://supervisord.org/) on Unix-like systems, without replacing init. Run each program [in the foreground, not as a daemon](https://supervisord.org/subprocess.html#nondaemonizing-of-subprocesses), so Supervisor can control it. + +Flower is a web app showing the status of Celery workers and tasks in real time, and it's [Celery's recommended monitor](https://docs.celeryq.dev/en/stable/userguide/monitoring.html#flower-real-time-celery-web-monitor). [Start it with `celery -A flower`](https://flower.readthedocs.io/en/latest/install.html). + +BorgBackup makes compressed, deduplicated backups [with authenticated encryption](https://www.borgbackup.org/), so your backup server only ever sees ciphertext. [Keep a copy of your key](https://borgbackup.readthedocs.io/en/stable/quickstart.html#repository-encryption) with `borg key export`. + +Chaos Toolkit runs chaos engineering experiments. [Each experiment declares](https://chaostoolkit.org/reference/concepts/) a steady-state hypothesis, which describes what normal looks like for your system. Its method runs actions and probes, and rollbacks can revert those actions. + +Whatever cloud your code talks to, keep access keys out of it. On the cloud's own machines, use the identity the machine already has: [an IAM role on EC2](https://docs.aws.amazon.com/boto3/latest/guide/credentials.html#best-practices-for-configuring-credentials), [a managed identity on Azure](https://learn.microsoft.com/en-us/azure/developer/python/sdk/authentication/overview), [the attached service account on Google Cloud](https://docs.cloud.google.com/docs/authentication/application-default-credentials#attached-sa). On your own machine, sign in with the cloud's command-line tool. The SDK's [credential chain](https://learn.microsoft.com/en-us/azure/developer/python/sdk/authentication/credential-chains) picks up that login, so the same code runs in both places. diff --git a/website/data/category_intros/distributed-computing.md b/website/data/category_intros/distributed-computing.md new file mode 100644 index 00000000..10bbf4f8 --- /dev/null +++ b/website/data/category_intros/distributed-computing.md @@ -0,0 +1,22 @@ +From a laptop to a cluster, Python distributed computing comes down to the workload: joblib runs loops, Dask scales pandas, Ray scales ML, and PySpark runs SQL. + +How to choose: + +- Parallel for loops on one machine: joblib +- Scaling pandas, NumPy, or scikit-learn code: Dask +- Training, tuning, and serving ML models, or GPU jobs: Ray +- Your own Python code with stateful workers on a cluster: Ray Core +- ETL and SQL on structured data, or a JVM shop: PySpark +- MPI programs on HPC clusters and supercomputers: mpi4py + +joblib is a [package for parallel computing and disk-based caching](https://joblib.readthedocs.io/en/stable/) that leaves your code as unmodified as possible. You [write a parallel for loop](https://joblib.readthedocs.io/en/stable/user_guide/parallel.html#common-usage) as a generator expression: `Parallel(n_jobs=2)(delayed(sqrt)(i ** 2) for i in range(10))`. By default, joblib runs the calls in separate worker processes. When one machine isn't enough, [switch its backend](https://joblib.readthedocs.io/en/stable/user_guide/parallel.html#setting-up-joblib-s-backend-with-parallel-config), and the same loop runs on a Dask, Ray, or Spark cluster. scikit-learn's [`n_jobs` runs on joblib](https://scikit-learn.org/stable/computing/parallelism.html) too, so it follows the backend you pick. + +Dask [scales pandas, scikit-learn, and NumPy workflows](https://docs.dask.org/en/stable/why.html) with minimal rewriting. A Dask DataFrame is [a collection of pandas dataframes](https://docs.dask.org/en/stable/) on different computers, and Dask Arrays parallelize NumPy. Dask [runs without any setup](https://docs.dask.org/en/stable/deploying.html#local-machine) on your laptop. Its `LocalCluster` follows the same interface as every other Dask cluster manager, so you swap it out when you're ready to scale up. + +Ray is a [unified framework for scaling AI and Python applications](https://docs.ray.io/en/latest/ray-overview/index.html). Its libraries each distribute one ML task: Ray Data, Train, Tune, Serve, and RLlib. Ray Core runs your own code: [decorate a function with `@ray.remote`](https://docs.ray.io/en/latest/ray-core/walkthrough.html), call it with `.remote()`, and fetch the result with `ray.get()`. For workers that keep state between calls, decorate a class the same way to get an actor. Ray runs on one machine with `ray.init()`, and on several nodes once you [deploy a Ray cluster](https://docs.ray.io/en/latest/cluster/getting-started.html). Ray Data [suits GPU workloads for deep learning inference](https://docs.ray.io/en/latest/data/comparisons.html#how-does-ray-data-compare-to-other-solutions-for-offline-inference) better than Spark does, but unlike Spark, it has no SQL interface. + +PySpark is the Python API for Apache Spark, for large-scale data processing. Dask's own comparison suggests Spark [when you prefer SQL, run mostly JVM infrastructure, or want an all-in-one solution](https://docs.dask.org/en/stable/spark.html#reasons-you-might-choose-spark). PySpark's docs [recommend DataFrames over RDDs](https://spark.apache.org/docs/latest/api/python/index.html), so Spark builds the most efficient query for you. Write it in [SQL or the DataFrame API](https://spark.apache.org/docs/latest/api/python/user_guide/sql.html), whichever you think in, and switch between the two as you go. Installing PySpark with pip is [for local use or as a client](https://spark.apache.org/docs/latest/api/python/getting_started/install.html) that connects to a cluster, not for setting up the cluster itself. The `pyspark` package also needs Java. + +mpi4py provides [Python bindings for MPI](https://mpi4py.readthedocs.io/en/stable/), the Message Passing Interface, so your Python code runs across workstations, clusters, and supercomputers. Lowercase methods like `comm.send` pass any picklable Python object, and uppercase ones like `comm.Send` [pass NumPy arrays the fast way](https://mpi4py.readthedocs.io/en/stable/tutorial.html). Run your script with `mpiexec -n 4 python -m mpi4py script.py`, so an unhandled exception [aborts the whole MPI run instead of deadlocking](https://mpi4py.readthedocs.io/en/stable/mpi4py.run.html#exceptions-and-deadlocks). mpi4py runs on an MPI implementation like MPICH or Open MPI, and in production its docs [recommend a custom-built or system-provided one](https://mpi4py.readthedocs.io/en/stable/install.html). + +Start on one machine: parallelism [brings extra complexity and overhead](https://docs.dask.org/en/stable/best-practices.html#start-small), and often you don't need it. When you do scale out, collect results at the end, since Dask's `compute()` and Ray's `ray.get()` both block until the work finishes. Call [`compute()` once](https://docs.dask.org/en/stable/best-practices.html#avoid-calling-compute-repeatedly) rather than in a loop, and [`ray.get()` as late as possible](https://docs.ray.io/en/latest/ray-core/tips-for-first-time.html#tip-1-delay-ray-get). diff --git a/website/data/category_intros/distribution.md b/website/data/category_intros/distribution.md new file mode 100644 index 00000000..531e32ac --- /dev/null +++ b/website/data/category_intros/distribution.md @@ -0,0 +1,25 @@ +Shipping to users without Python means converting a Python script to an executable, which PyInstaller does in one command. Pyarmor can obfuscate it first. + +How to choose: + +- Standalone executable for most apps: PyInstaller +- Obfuscated scripts, bound to a machine or set to expire: Pyarmor +- Compiled code instead of bundled bytecode: Nuitka +- Tools for machines that already run Python: shiv +- Installers and Linux packages: cx_Freeze + +PyInstaller [bundles your app and all its dependencies](https://pyinstaller.org/en/latest/) into one package, so users run it without installing Python or any modules. For most programs, that's [one short command](https://pyinstaller.org/en/latest/operating-mode.html): `pyinstaller myscript.py`. + +Pyarmor [obfuscates Python scripts](https://pyarmor.readthedocs.io/en/latest/tutorial/getting-started.html), and can bind them to a machine or make them expire. `pyarmor gen foo.py` writes the obfuscated script to `dist/`. To ship an executable, Pyarmor [packs through PyInstaller](https://pyarmor.readthedocs.io/en/latest/tutorial/obfuscation.html#packing-obfuscated-scripts): `pyarmor gen --pack onefile foo.py`. + +Without obfuscation, a PyInstaller bundle holds `.pyc` files that [could in principle be decompiled](https://pyinstaller.org/en/latest/operating-mode.html#hiding-the-source-code). Pyarmor is commercial: its free version is only for scripts that [won't make you a lot of money](https://pyarmor.readthedocs.io/en/latest/licenses.html#terms-of-use). + +Nuitka is an optimizing Python compiler, and it [needs a C compiler](https://nuitka.net/user-documentation/user-manual.html). Its default mode needs Python on the machine, so [build in standalone mode to distribute](https://nuitka.net/user-documentation/tutorial-setup-and-build.html#distribute), and copy the resulting folder. Compiling protects your source code, but Nuitka's docs say constants stay readable unless you buy [Nuitka Commercial](https://nuitka.net/doc/commercial.html). + +shiv builds [self-contained zipapps with all their dependencies included](https://shiv.readthedocs.io/en/latest/). `shiv -c hello -o hello .` packages your project the way `pip install .` would, with `-c` naming its console script. The result [depends on a pre-installed Python](https://packaging.python.org/en/latest/overview/#depending-on-a-pre-installed-python), which you can count on in your data centers and on developers' machines, but not on every user's computer. + +cx_Freeze [freezes a script and its modules into a standalone executable](https://cx-freeze.readthedocs.io/en/latest/script.html), and also [builds installers and packages](https://cx-freeze.readthedocs.io/en/latest/): MSI for Windows, DMG for macOS, and deb, RPM, and AppImage for Linux. Put your options in `pyproject.toml` under `[tool.cxfreeze]` and build with [`cxfreeze build`](https://cx-freeze.readthedocs.io/en/latest/setup_script.html). + +Get a folder build working before you switch to a single file, since problems are easier to diagnose in a folder. PyInstaller's docs [say so for one-folder mode](https://pyinstaller.org/en/latest/operating-mode.html#bundling-to-one-file), and Nuitka's [for standalone mode](https://nuitka.net/user-documentation/use-cases.html#standalone-program-distribution). + +For an executable, plan to build on each OS you ship to. PyInstaller [isn't a cross-compiler](https://pyinstaller.org/en/latest/), cx_Freeze [only makes executables for the platform it runs on](https://cx-freeze.readthedocs.io/en/latest/faq.html#freezing-for-other-platforms), and Nuitka's docs suggest [Nuitka-Action](https://nuitka.net/user-documentation/use-cases.html#building-with-github-workflows) to build on all 3 OSes in GitHub workflows. diff --git a/website/data/category_intros/documentation.md b/website/data/category_intros/documentation.md new file mode 100644 index 00000000..c0fad335 --- /dev/null +++ b/website/data/category_intros/documentation.md @@ -0,0 +1,21 @@ +Write only docstrings, and pdoc is all the Python documentation generator you need. Write guides too, and build the site with Sphinx or Material for MkDocs. + +How to choose: + +- API docs straight from docstrings, with no configuration: pdoc +- Handwritten docs plus an API reference, as HTML, PDF, and more: Sphinx +- A searchable Markdown site, with no HTML, CSS, or JavaScript to learn: Material for MkDocs +- Architecture diagrams as code, kept in version control: Diagrams +- A Material for MkDocs site, new or existing, on its team's own generator: Zensical + +pdoc [aims to do one thing and do it well](https://pdoc.dev/docs/pdoc.html#what-is-pdoc): API documentation that follows your module hierarchy, with no configuration. Docstrings are Markdown, and it understands Google and numpydoc styles too. Run `pdoc ./demo.py` or `pdoc my_module_name` for a [preview in your browser that reloads](https://pdoc.dev/docs/pdoc.html#quickstart) when you edit the code, then `pdoc ./demo.py -o ./docs` to export the HTML. Its output is self-contained HTML, and for substantially more complex documentation needs, [pdoc's docs recommend Sphinx](https://pdoc.dev/docs/pdoc.html#limitations). + +Sphinx [focuses on handwritten documentation](https://www.sphinx-doc.org/en/stable/usage/quickstart.html) and turns one set of source files into HTML, a PDF via LaTeX, man pages, and more. Its default markup is reStructuredText, and it can [read Markdown through MyST-Parser](https://www.sphinx-doc.org/en/stable/usage/markdown.html). Run `sphinx-quickstart` to set up a source directory with a `conf.py`, then list your pages in the root document's toctree. `make html` builds the site, and `make latexpdf` the PDF. To document your code, [autodoc](https://www.sphinx-doc.org/en/stable/usage/extensions/autodoc.html) pulls in its docstrings, which you mix with your handwritten pages. + +Material for MkDocs is [a documentation framework on top of MkDocs](https://squidfunk.github.io/mkdocs-material/getting-started/), and `pip install mkdocs-material` installs MkDocs with it. You write Markdown and get a searchable static site, with [no HTML, CSS, or JavaScript to know](https://squidfunk.github.io/mkdocs-material/). Run `mkdocs new .`, then [set `site_name`, `site_url`, and `theme: name: material`](https://squidfunk.github.io/mkdocs-material/creating-your-site/#minimal-configuration) in `mkdocs.yml`, and `mkdocs serve` previews the site as you write. For reference docs from docstrings, try [mkdocstrings](https://squidfunk.github.io/mkdocs-material/alternatives/#sphinx) before switching to Sphinx: Material's alternatives page says it builds on MkDocs and adds Sphinx-like functionality. + +Zensical is [a static site generator from the creators of Material for MkDocs](https://zensical.org/docs/get-started/), with the same batteries-included approach. After `pip install zensical`, [`zensical new .`](https://zensical.org/docs/create-your-site/) creates a `docs/` folder, a `zensical.toml` config, and a GitHub Actions workflow. `zensical serve` previews as you write, and `zensical build` writes the static site. It also [builds existing MkDocs projects without changes](https://zensical.org/docs/compatibility/mkdocs/) from their `mkdocs.yml`, and its classic theme variant keeps the Material for MkDocs look. + +Diagrams [draws cloud system architecture in Python code](https://diagrams.mingrammer.com/), so you can track changes to a diagram in version control. It renders with Graphviz, so [install Graphviz first](https://diagrams.mingrammer.com/docs/getting-started/installation), then `pip install diagrams`. Describe the system in a `with Diagram("Web Service", show=False):` block and chain nodes with `>>`. Running `python diagram.py` saves it as a PNG in your working directory. + +Install [Sphinx](https://www.sphinx-doc.org/en/stable/usage/installation.html), [Material for MkDocs](https://squidfunk.github.io/mkdocs-material/getting-started/#with-pip), or [Zensical](https://zensical.org/docs/get-started/#install-with-pip) into your project's virtual environment, as each one's docs recommend. The documentation generators here all write static HTML, so publish it from CI. All four document a GitHub Actions workflow that deploys to GitHub Pages: [Sphinx](https://www.sphinx-doc.org/en/stable/tutorial/deploying.html#publishing-your-html-documentation), [Material for MkDocs](https://squidfunk.github.io/mkdocs-material/publishing-your-site/), [Zensical](https://zensical.org/docs/publish-your-site/), and [pdoc](https://pdoc.dev/docs/pdoc.html#deploying-to-github-pages). diff --git a/website/data/category_intros/email.md b/website/data/category_intros/email.md new file mode 100644 index 00000000..b7b78a6c --- /dev/null +++ b/website/data/category_intros/email.md @@ -0,0 +1,11 @@ +One function call sends mail through Gmail or another SMTP server with yagmail. This Python email library aims to make sending mail painless. + +How to choose: + +- Gmail: yagmail, with an app password or OAuth2 +- Another SMTP server: yagmail, pointed at its host +- An asyncio app: yagmail's async client + +yagmail is a [wrapper around smtplib's SMTP connection](https://yagmail.readthedocs.io/en/latest/api.html#yagmail.Client) that connects to Gmail unless you pass another `host`. It builds the message for you: call `send()` with the recipients, a subject, and `contents`, a list whose strings it [reads as a local file, HTML, or text](https://yagmail.readthedocs.io/en/latest/usage.html#magical-contents). So one call sends text, HTML, and attachments. Wrap a string in `yagmail.raw` when it must stay plain text. In asyncio code, use its async client [as an async context manager](https://yagmail.readthedocs.io/en/latest/usage.html#starting-and-closing-connections). + +Keep your password out of your script. Install `yagmail[all]` to get keyring, then [register your credentials once](https://yagmail.readthedocs.io/en/latest/setup.html#configuring-credentials) with `yagmail.register()`, and yagmail reads them from your system keyring. For Gmail, that password is an app password. For credentials you can revoke, [use OAuth2](https://yagmail.readthedocs.io/en/latest/setup.html#using-oauth2): whoever gets its token file can send mail, but nothing else. diff --git a/website/data/category_intros/environment-management.md b/website/data/category_intros/environment-management.md new file mode 100644 index 00000000..599a379c --- /dev/null +++ b/website/data/category_intros/environment-management.md @@ -0,0 +1,15 @@ +Since uv both downloads Python and creates virtual environments, one Python environment manager does the jobs that pyenv and virtualenv split between them. + +How to choose: + +- Python versions and virtual environments from one tool: uv +- A Python API for creating environments, or plugins: virtualenv +- Switching the `python` command between versions, and nothing more: pyenv + +uv [manages Python versions](https://docs.astral.sh/uv/) as well as packages. When a command needs a Python you don't have, uv [downloads it for you](https://docs.astral.sh/uv/guides/install-python/), so you don't need Python installed to get started. A project's virtual environment lives in a `.venv` folder [inside the project, where editors can find it](https://docs.astral.sh/uv/concepts/projects/layout/#the-project-environment); keep it out of version control. Outside a project, `uv venv --python ` [creates an environment](https://docs.astral.sh/uv/pip/environments/) with that version, and downloads it if needed. + +virtualenv [creates isolated Python environments](https://virtualenv.pypa.io/en/latest/), and a subset of it ships with Python as the venv module. Its docs place it [between venv and uv](https://virtualenv.pypa.io/en/latest/explanation.html#virtualenv-vs-venv-vs-uv): faster and more featureful than venv, and still pure Python. Pick it when you need plugins, or a Python API to create environments. It uses the Python it runs under unless you [pass `-p`](https://virtualenv.pypa.io/en/latest/how-to/usage.html#select-a-python-version) with another installed version. As a command-line tool, it [belongs in an isolated environment](https://virtualenv.pypa.io/en/latest/how-to/install.html), not your system Python. + +pyenv does one job: it [lets you switch between multiple versions of Python](https://github.com/pyenv/pyenv), and leaves virtual environments [to you or its pyenv-virtualenv plugin](https://github.com/pyenv/pyenv#in-contrast-with-pythonbrew-and-pythonz-pyenv-does-not). Install it with the [automatic installer](https://github.com/pyenv/pyenv#1-automatic-installer-recommended) or, on macOS, Homebrew, then [set up your shell](https://github.com/pyenv/pyenv#b-set-up-your-shell-environment-for-pyenv) for it. [Choose the version](https://github.com/pyenv/pyenv#switch-between-python-versions) with `pyenv shell` for the session, `pyenv local` for a directory, or `pyenv global` for your user account. [Most versions are built from source](https://github.com/pyenv/pyenv#install-additional-python-versions) as you install them, so install Python's build dependencies first. pyenv [doesn't work on Windows](https://github.com/pyenv/pyenv#windows) outside the Windows Subsystem for Linux, so on Windows, use uv, which [supports Windows](https://docs.astral.sh/uv/). + +Whichever tools you combine, name each project's Python in a `.python-version` file. `pyenv local` [writes that file](https://github.com/pyenv/pyenv#understanding-python-version-selection), and so does `uv python pin`, whose docs [recommend a plain version number](https://docs.astral.sh/uv/concepts/python-versions/#python-version-files) there so other tools can read it. virtualenv [uses the version pyenv selected](https://virtualenv.pypa.io/en/latest/how-to/usage.html#using-version-managers-pyenv-mise-asdf) with no extra configuration. Don't install packages into the Python that came with your operating system, which [often manages its packages itself](https://docs.astral.sh/uv/pip/environments/); a virtual environment keeps yours apart. diff --git a/website/data/category_intros/erp.md b/website/data/category_intros/erp.md new file mode 100644 index 00000000..390c2921 --- /dev/null +++ b/website/data/category_intros/erp.md @@ -0,0 +1,19 @@ +Out of the box, Odoo's CRM, eCommerce, and accounting apps run alone or combine into a full Python ERP. Anything they lack, you add as a module. + +How to choose: + +- Trying Odoo, or customizing it without code: Odoo Online +- Writing your own modules: a source install +- Hosting your own modules in the cloud: Odoo.sh +- Free and open source: Odoo Community +- More features, with support and upgrades: Odoo Enterprise + +Everything in Odoo starts and ends with modules, and [the main user-facing ones are flagged as Apps](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/01_architecture.html#odoo-modules). Odoo Enterprise is [extra modules installed on top of Community](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/01_architecture.html#odoo-editions). [Community is free and open source under the LGPL](https://www.odoo.com/documentation/latest/administration.html#editions). Enterprise is shared source, and its license [ties running it to an Enterprise subscription](https://www.odoo.com/documentation/latest/legal/licenses.html). + +[Odoo Online](https://www.odoo.com/documentation/latest/administration/odoo_online.html) runs in your browser with nothing to install, and handles customizations that need no code. Your own modules need Odoo.sh or your own server. Odoo.sh is Odoo's official cloud platform, and it [builds your modules from a GitHub repository](https://www.odoo.com/documentation/latest/administration/odoo_sh/create_module.html). + +To develop modules, Odoo's developer docs prefer [a source install](https://www.odoo.com/documentation/latest/developer/tutorials/setup_guide.html), which runs Odoo straight from its code. Odoo's logic is [written in Python, and it stores data only in PostgreSQL](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/01_architecture.html#multitier-application). Keep your modules in a directory of their own, and [start the server with `odoo-bin`](https://www.odoo.com/documentation/latest/administration/on_premise/source.html#running-odoo), adding that directory to `--addons-path`. Then work through [Server framework 101](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101.html), which builds one module chapter by chapter. + +A module can [add new business logic or change what's already there](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/01_architecture.html#odoo-modules). To change a standard model, extend it from your own module: [model inheritance](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/12_inheritance.html#model-inheritance) adds fields and overrides methods on a model another module defines. Screens work the same way: [view inheritance](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/12_inheritance.html#view-inheritance) applies your extension views on top of the originals instead of overwriting them. + +Before you write a module, [check whether Odoo already covers the case](https://www.odoo.com/documentation/latest/developer/tutorials/server_framework_101/02_newapp.html). A database with custom modules [can't be upgraded until they're ready for the new version](https://www.odoo.com/documentation/latest/administration/upgrade.html), so [cut what duplicates the standard modules](https://www.odoo.com/documentation/latest/developer/howtos/upgrade_custom_db.html#step-1-stop-the-developments). diff --git a/website/data/category_intros/file-format-processing.md b/website/data/category_intros/file-format-processing.md new file mode 100644 index 00000000..069f0bec --- /dev/null +++ b/website/data/category_intros/file-format-processing.md @@ -0,0 +1,51 @@ +PDF, Word, and Excel files open with pypdf, python-docx, and openpyxl, a Python file format library for each. MarkItDown turns all three into Markdown. + +How to choose: + +- Splitting, merging, and reading PDFs: pypdf +- Word documents: python-docx +- Excel: openpyxl to read or edit, XlsxWriter for new reports +- Documents to Markdown for an LLM: MarkItDown +- ELF binaries and DWARF debug info: pyelftools +- One table exported to CSV, JSON, Excel, and more: Tablib +- Scanned PDFs, tables, and complex layouts: Docling +- PowerPoint decks: python-pptx +- New PDFs drawn from Python code: ReportLab +- PDF text with its position and font: pdfminer.six +- PDFs from HTML and CSS: WeasyPrint +- Markdown to HTML: markdown-it-py for CommonMark, Python-Markdown for its extensions, Mistune for speed +- Config files: tomllib for TOML, PyYAML for YAML + +pypdf works on PDFs that already exist: it [splits, merges, crops, and transforms pages](https://pypdf.readthedocs.io/en/stable/), adds passwords, and pulls out text and metadata. It's [pure Python with no C dependency](https://pypdf.readthedocs.io/en/stable/meta/comparisons.html), and it doesn't create PDFs. pypdf [isn't OCR software](https://pypdf.readthedocs.io/en/stable/user/extract-text.html), so run scanned pages through OCR instead. + +python-docx [only edits existing documents](https://python-docx.readthedocs.io/en/latest/user/documents.html): `Document()` opens a built-in template with no content. Start from your own .docx instead, so its styles, headers, and footers carry over. + +openpyxl reads and writes Excel files. For big workbooks, open them with `read_only=True` or create them with `write_only=True`, which [keep memory near constant](https://openpyxl.readthedocs.io/en/stable/optimized.html). A cell with a formula loads the formula; pass [`data_only=True`](https://openpyxl.readthedocs.io/en/stable/tutorial.html) to get the value Excel last calculated. + +XlsxWriter [only writes new files](https://xlsxwriter.readthedocs.io/introduction.html) and can't read or modify existing ones, but it supports more Excel features than the alternatives. Open the workbook [in a `with` block](https://xlsxwriter.readthedocs.io/workbook.html) so it gets closed and saved. For large files, turn on [`constant_memory`](https://xlsxwriter.readthedocs.io/working_with_memory.html) and write the rows in order. From pandas, pass [`engine='xlsxwriter'`](https://xlsxwriter.readthedocs.io/working_with_pandas.html) to `pd.ExcelWriter`. + +MarkItDown converts files to Markdown [for LLMs and text analysis](https://github.com/microsoft/markitdown), not for high-fidelity conversions that people read. Install `markitdown[all]`, or only the extras for your formats, like `markitdown[pdf, docx, pptx]`. + +Docling [understands PDF layout](https://docling-project.github.io/docling/): reading order, tables, and formulas. It also runs OCR on scanned pages. Its models [run locally and send no data out](https://docling-project.github.io/docling/usage/advanced_options/#using-remote-services) unless you turn remote services on. Each file becomes a DoclingDocument, which you export to Markdown or [split into chunks](https://docling-project.github.io/docling/concepts/chunking/) for an embedding model. + +python-pptx builds decks from data, like a database query or analytics output, and [doesn't need PowerPoint installed](https://python-pptx.readthedocs.io/en/latest/). It [only edits existing presentations](https://python-pptx.readthedocs.io/en/latest/user/presentations.html), so start from your own deck: its theme, slide master, and slide layouts set how the slides look. Add each slide from one of those layouts, picked by [its index in your deck](https://python-pptx.readthedocs.io/en/latest/user/slides.html). + +ReportLab draws new PDFs from Python code. Learn it on [`pdfgen`, its lowest-level interface](https://docs.reportlab.com/developerfaqs/), then build multi-page documents with [Platypus](https://docs.reportlab.com/reportlab/userguide/ch5_platypus/). Platypus lets you keep paragraph styles and page layouts in one shared file, so restyling takes a few lines. + +pdfminer.six [focuses on text](https://github.com/pdfminer/pdfminer.six): it gets each piece of text with its exact location, font, and color. Start with `extract_text()` from its [high-level API](https://pdfminersix.readthedocs.io/en/latest/tutorial/highlevel.html). A PDF [stores only characters and their positions](https://pdfminersix.readthedocs.io/en/latest/topic/converting_pdf_to_text.html), so pdfminer.six guesses words, lines, and paragraphs from the layout. Tune those guesses with `LAParams`. + +WeasyPrint turns HTML and CSS into PDFs, like [reports, invoices, and tickets](https://doc.courtbouillon.org/weasyprint/stable/). It runs [no JavaScript](https://doc.courtbouillon.org/weasyprint/stable/going_further.html). Set page size and margins [with the CSS `@page` rule](https://doc.courtbouillon.org/weasyprint/stable/common_use_cases.html). + +markdown-it-py [follows the CommonMark spec](https://markdown-it-py.readthedocs.io/en/latest/) and takes plugins for more syntax. For content your users submit, use the [`js-default` preset](https://markdown-it-py.readthedocs.io/en/latest/security.html), since the default settings aren't safe for it. + +Python-Markdown [isn't a CommonMark implementation](https://python-markdown.github.io/): it follows the original Markdown syntax and has an extension API. It [doesn't sanitize its HTML output](https://python-markdown.github.io/sanitization/), so sanitize it yourself when the input is untrusted. + +Mistune is [fast and has no dependencies](https://mistune.lepture.com/en/latest/). For untrusted input, build the parser with [`mistune.create_markdown()`](https://mistune.lepture.com/en/latest/guide.html), which escapes HTML tags, since `mistune.html()` doesn't. + +tomllib [only reads TOML](https://docs.python.org/3/library/tomllib.html), from a file opened in binary mode. + +Tablib holds one dataset and exports it to many formats; Excel, YAML, and pandas [are optional extras](https://tablib.readthedocs.io/en/stable/formats.html), like `tablib[xlsx]`. + +pyelftools is [pure Python with no dependencies](https://github.com/eliben/pyelftools); start from its [`ELFFile` class](https://github.com/eliben/pyelftools/blob/main/doc/user-guide.md) and stay on the high-level API. + +Treat every file you didn't create as untrusted. With PyYAML, call [`yaml.safe_load()`](https://pyyaml.org/wiki/PyYAMLDocumentation#loading-yaml), never `yaml.load()`, which can run any Python function. Install [defusedxml](https://openpyxl.readthedocs.io/en/stable/#security) next to openpyxl to guard against XML attacks like billion laughs. [Catch pypdf's exceptions](https://pypdf.readthedocs.io/en/stable/user/security.html) yourself, so a broken PDF can't crash your service. For MarkItDown, call [`convert_local()` or `convert_stream()`](https://github.com/microsoft/markitdown#security-considerations) instead of `convert()`, which also fetches remote URIs. Cap Docling's input with [`max_num_pages` and `max_file_size`](https://docling-project.github.io/docling/usage/advanced_options/#impose-limits-on-the-document-size). Run WeasyPrint on untrusted HTML [as a user with limited access](https://doc.courtbouillon.org/weasyprint/stable/first_steps.html#security), with a URL fetcher that blocks local files. With tomllib, [limit the size](https://docs.python.org/3/library/tomllib.html) of the data you parse. diff --git a/website/data/category_intros/file-manipulation.md b/website/data/category_intros/file-manipulation.md new file mode 100644 index 00000000..ea7aeb0c --- /dev/null +++ b/website/data/category_intros/file-manipulation.md @@ -0,0 +1,19 @@ +Without installing anything, two Python file manipulation libraries cover paths and file types by name: pathlib and mimetypes. python-magic reads the bytes. + +How to choose: + +- Building and joining paths, and listing directories: pathlib +- A MIME type from a file name or URL: mimetypes +- A file's type from its content, like the Unix `file` command: python-magic +- Rerunning or reloading code when files change: watchfiles +- Your own handler for each created, modified, or moved file: watchdog + +If you aren't sure which pathlib class you need, the docs say [`Path` is most likely it](https://docs.python.org/3/library/pathlib.html): it makes a concrete path for the platform your code runs on. [Join paths with the `/` operator](https://docs.python.org/3/library/pathlib.html#basic-use), and list a directory with `iterdir()` or `glob()`. + +mimetypes [maps a file name's extension to a MIME type](https://docs.python.org/3/library/mimetypes.html), and a MIME type back to extensions. A type guess returns 2 values: a type for the Content-Type header, and an encoding like gzip for the Content-Encoding header. The type is [`None` for a missing or unknown suffix](https://docs.python.org/3/library/mimetypes.html#mimetypes.guess_type). + +python-magic wraps libmagic, which [identifies file types by checking their headers](https://github.com/ahupp/python-magic), the way the Unix `file` command does. It's a thin wrapper, so install libmagic too, with `apt-get install libmagic1` or `brew install libmagic`. Call `magic.from_file(path, mime=True)` for a MIME type, or `magic.from_buffer()` on bytes you already have. + +watchfiles is built for [file watching and code reload](https://watchfiles.helpmanual.io/), and its Rust core groups changes into batches instead of firing once per file. `watch()` is a generator that [yields sets of changes](https://watchfiles.helpmanual.io/api/watch/#watchfiles.watch), and `awatch()` is the async version. To restart code when files change, use [`run_process()`](https://watchfiles.helpmanual.io/api/run_process/#watchfiles.run_process) with a function or a command. From a shell, the [`watchfiles` CLI](https://watchfiles.helpmanual.io/cli/) does the same: `watchfiles --filter python 'pytest --lf' src tests`. + +watchdog gives you [an API and a shell tool](https://python-watchdog.readthedocs.io/en/latest/) for monitoring directories. Its [quickstart](https://python-watchdog.readthedocs.io/en/latest/quickstart.html) has you subclass `FileSystemEventHandler`, override methods like `on_created()` and `on_modified()`, schedule the handler on an `Observer`, and start that thread. An observer skips subdirectories unless you pass `recursive=True` to `schedule()`. diff --git a/website/data/category_intros/functional-programming.md b/website/data/category_intros/functional-programming.md new file mode 100644 index 00000000..6374a009 --- /dev/null +++ b/website/data/category_intros/functional-programming.md @@ -0,0 +1,21 @@ +Beyond functools, install more-itertools, as the itertools docs suggest. toolz is a fuller Python functional programming library, and returns adds typed errors. + +How to choose: + +- Partial application, decorators, and caching: functools +- More iterator tools, the itertools recipes included: more-itertools +- Composing functions into pipelines, currying, and dict helpers: toolz, or cytoolz for speed +- Everyday helpers for collections, decorators, retries, and debugging: funcy +- Errors and missing values as typed containers checked by mypy: returns + +functools is the standard library's module [for higher-order functions](https://docs.python.org/3/library/functools.html), functions that act on or return other functions. Python's Functional Programming HOWTO calls `partial()` [the most useful tool in the module](https://docs.python.org/3/howto/functional.html#the-functools-module): it fills in some of a function's arguments and gives you a new function. When you write a decorator, wrap its inner function with [`wraps`](https://docs.python.org/3/library/functools.html#functools.wraps), so the decorated function keeps its name and docstring. The same HOWTO finds many uses of `reduce()` [clearer as a `for` loop](https://docs.python.org/3/howto/functional.html#small-functions-and-the-lambda-expression). + +The itertools docs point you to more-itertools for [their recipes and many more](https://docs.python.org/3/library/itertools.html#itertools-recipes). It collects [building blocks beyond itertools](https://more-itertools.readthedocs.io/en/stable/), for grouping, windowing, lookahead, and more. The itertools recipes sit in its top-level package, so `from more_itertools import flatten` works. + +toolz [extends itertools and functools](https://toolz.readthedocs.io/en/latest/) with functions that are composable, pure, and lazy, and its API [follows Clojure's standard library](https://toolz.readthedocs.io/en/latest/heritage.html). Each function takes and returns only iterables, dictionaries, and functions, so they [compose to solve your own problems](https://toolz.readthedocs.io/en/latest/composition.html). [`pipe`](https://toolz.readthedocs.io/en/latest/api.html#toolz.functoolz.pipe) runs a value through a sequence of functions, like pipes in Unix. Stick with `partial` at first, and once it shows up several times in your code, [switch to the `toolz.curried` namespace](https://toolz.readthedocs.io/en/latest/curry.html#curry). toolz is a general-purpose library, and for data analytics its docs say [a library built for it](https://toolz.readthedocs.io/en/latest/streaming-analytics.html#disclaimer) may serve you better. [cytoolz](https://github.com/pytoolz/cytoolz) implements the same API in Cython, as a drop-in replacement when you need more speed. + +funcy is a collection of functional tools [focused on practicality](https://github.com/Suor/funcy), inspired by Clojure and underscore. Next to sequence tools, it has [collection functions that keep the type](https://funcy.readthedocs.io/en/stable/overview.html) of a dict or set. It also has control flow helpers, like `@retry` and `silent`, and debugging helpers, like `tap` and `log_calls`. Many of its functions take a regex, a mapping, or a set [where you'd pass a function](https://funcy.readthedocs.io/en/stable/extended_fns.html#extended-function-semantics). + +returns puts results in typed containers: [`Maybe` for None and `Result` for exceptions](https://returns.readthedocs.io/en/latest/pages/quickstart.html#why), plus `IO` for impure code and `Future` for async code. Its docs [really recommend mypy](https://returns.readthedocs.io/en/latest/pages/quickstart.html#typechecking-and-other-integrations), and typing [only works correctly with its mypy plugin](https://returns.readthedocs.io/en/latest/pages/result.html). So it fits projects that check types with mypy. Turn functions that raise into ones that return a `Result` with [`@safe`](https://returns.readthedocs.io/en/latest/pages/result.html#safe), and chain the steps with [`flow`](https://returns.readthedocs.io/en/latest/pages/pipeline.html#flow), which its docs call the recommended way to write code with returns. + +Mix functional style with the rest of your code: Python's HOWTO says functional-style programs usually [give a functional-appearing interface](https://docs.python.org/3/howto/functional.html) and use non-functional features inside. For a plain map or filter, toolz's own docs call comprehensions [more Pythonic](https://toolz.readthedocs.io/en/latest/streaming-analytics.html). Learn a core set of functions; toolz says [about a dozen covers most tasks](https://toolz.readthedocs.io/en/latest/control.html), and the right word only helps when your readers know it too. diff --git a/website/data/category_intros/game-development.md b/website/data/category_intros/game-development.md new file mode 100644 index 00000000..60c4a8c2 --- /dev/null +++ b/website/data/category_intros/game-development.md @@ -0,0 +1,21 @@ +A 2D game needs a loop: write your own with pygame-ce, the Python game development library, or let Arcade run it. Panda3D does 3D, and Ren'Py visual novels. + +How to choose: + +- 2D game where you write the game loop: pygame-ce +- 2D game with a ready-made loop and physics engines: Arcade +- 3D game: Panda3D +- Visual novel or life simulation game: Ren'Py +- Windowing, OpenGL graphics, and sound with no other dependencies: pyglet + +pygame-ce is the community edition of pygame, a fork by its former core developers that aims for more frequent releases. Your code still says `import pygame`, so if pygame is already installed, [uninstall it first](https://github.com/pygame-community/pygame-ce/wiki/Installing-pygame%E2%80%90ce), then run `pip install pygame-ce` in a virtual environment. Its [quick start](https://pyga.me/docs/) gives you full control of the game loop: handle events, draw the frame, `flip()` the display, and cap the frame rate with `clock.tick()`. Convert each image once after you load it, with [`convert()`, or `convert_alpha()` if it has transparency](https://pyga.me/docs/ref/surface.html#pygame.Surface.convert), so it blits fast. It's under the LGPL, and its README says [closed-source and commercial games are fine](https://github.com/pygame-community/pygame-ce). + +Arcade is [an easy-to-learn library for 2D games](https://github.com/pythonarcade/arcade), built on pyglet and OpenGL, and meant for beginning programmers too. Start with the [Platformer Tutorial](https://api.arcade.academy/en/stable/tutorials/platform_tutorial/step_01.html): you subclass `arcade.Window`, draw in `on_draw()`, and `arcade.run()` runs the loop until the window closes. For movement and collisions, it comes with [physics engines](https://api.arcade.academy/en/stable/api_docs/api/physics_engines.html) for top-down and platformer games. Its code is MIT, and its [built-in assets need no attribution](https://api.arcade.academy/en/stable/), so you can ship them in a commercial game. + +Panda3D is [a 3D engine written in C++ with Python bindings](https://docs.panda3d.org/latest/python/introduction/index), and its manual says it's a tool for skilled programmers, not beginners. Install it with [`pip install panda3d`](https://github.com/panda3d/panda3d), subclass [`ShowBase`](https://docs.panda3d.org/latest/python/introduction/tutorial/starting-panda3d), and call `run()`, which holds the main loop. Its [`build_apps` tool](https://docs.panda3d.org/latest/python/distribution/index) builds self-contained executables for Windows, Linux, and macOS without needing each system. It's BSD, free for commercial games. + +Ren'Py is a [visual novel engine](https://www.renpy.org/) with its own script language, for stories that run on computers and mobile devices. It isn't a pip package: [download Ren'Py and run its launcher](https://www.renpy.org/doc/html/quickstart.html), create a project there, and write your story in `script.rpy`. Python works inside the scripts, and [third-party pure-Python packages](https://www.renpy.org/doc/html/python.html#first-and-third-party-python-modules-and-packages) go in `game/python-packages`. Ship with [Build Distributions](https://www.renpy.org/doc/html/build.html) in the launcher, which also builds a package for itch.io and Steam. Most of Ren'Py is MIT, but some parts are LGPL, so [distribute your game in a way that satisfies the LGPL](https://www.renpy.org/doc/html/license.html). + +pyglet is a [windowing and multimedia library with no external dependencies](https://pyglet.readthedocs.io/en/latest/), written in pure Python: windows, input, OpenGL graphics, images, video, and sound. Start with [Writing a pyglet application](https://pyglet.readthedocs.io/en/latest/programming_guide/quickstart.html), which attaches handlers with `@window.event` and calls `pyglet.app.run()`. Draw through a `Batch`, since the docs say [you always want batched rendering](https://pyglet.readthedocs.io/en/latest/programming_guide/shapes.html) for performance. It's under the BSD license. + +Move things by the time since the last frame, so your game runs at the same speed at any frame rate. pygame-ce's quick start gets it in seconds by [dividing `clock.tick()` by 1000](https://pyga.me/docs/), pyglet passes it as `dt` to [scheduled functions](https://pyglet.readthedocs.io/en/latest/programming_guide/time.html#sprite-movement-techniques), and Arcade passes it as `delta_time` to [`on_update()`](https://api.arcade.academy/en/stable/api_docs/api/window.html#arcade.Window.on_update). diff --git a/website/data/category_intros/geolocation.md b/website/data/category_intros/geolocation.md new file mode 100644 index 00000000..3c5757c3 --- /dev/null +++ b/website/data/category_intros/geolocation.md @@ -0,0 +1,20 @@ +Addresses become coordinates with geopy, whose Python geolocation API covers many geocoding services. GeoPandas analyzes map data, and GeoDjango serves it. + +How to choose: + +- Addresses to coordinates, and coordinates back to addresses: geopy +- The distance between two points: geopy +- Geographic data in tables and files, like shapefiles and GeoJSON: GeoPandas +- Web apps that store and query geographic data: GeoDjango +- A visitor's country or city from their IP address, in a Django project: GeoDjango +- Building, encoding, and validating GeoJSON objects by hand: geojson + +geopy is [a client for geocoding services, not a service itself](https://geopy.readthedocs.io/en/stable/#geopy-is-not-a-service), and each service has its own terms of use, quotas, and pricing. Every geocoder has a `geocode()` method that turns an address into a location, and most also have `reverse()` for the other way around. For OpenStreetMap's Nominatim, [set a `user_agent` that names your app](https://geopy.readthedocs.io/en/stable/#geopy.geocoders.Nominatim) and follow [its usage policy](https://operations.osmfoundation.org/policies/nominatim/). The policy bans heavy use, auto-complete, and systematic queries. It also asks you to cache results and show attribution, and it discourages bulk geocoding. To geocode a DataFrame, [wrap the call in `RateLimiter`](https://geopy.readthedocs.io/en/stable/#usage-with-pandas), which adds delays between requests and retries failed ones. Check first that your service allows bulk requests at all. For distances, `geopy.distance.distance` [computes the geodesic distance](https://geopy.readthedocs.io/en/stable/#module-geopy.distance) between two points. + +GeoPandas adds geometry columns to pandas, so you can [do in Python what would otherwise take a spatial database such as PostGIS](https://geopandas.org/en/stable/). Load a shapefile, GeoJSON, or GeoPackage with `read_file()`, which [detects the file type and returns a GeoDataFrame](https://geopandas.org/en/stable/getting_started/introduction.html), and write it back with `to_file()`. For a table of latitudes and longitudes, [build the points with `points_from_xy()`](https://geopandas.org/en/stable/gallery/create_geopandas_from_pandas.html) and set `crs="EPSG:4326"`. To geocode a column of addresses, GeoPandas [calls geopy for you](https://geopandas.org/en/stable/docs/user_guide/geocoding.html), so the geocoding service's terms still apply. + +GeoDjango is [a contrib module that ships with Django](https://docs.djangoproject.com/en/stable/ref/contrib/gis/tutorial/). It adds model fields for geometries, spatial queries to the ORM, and geometry editing to the admin. Its tutorial assumes you already know Django. Run it on PostGIS, which its docs recommend as [the most mature and feature-rich open source spatial database](https://docs.djangoproject.com/en/stable/ref/contrib/gis/install/#spatial-database). To load a shapefile, [`ogrinspect` writes the model and a `LayerMapping` imports the rows](https://docs.djangoproject.com/en/stable/ref/contrib/gis/tutorial/#importing-spatial-data). For IP geolocation, [GeoIP2](https://docs.djangoproject.com/en/stable/ref/contrib/gis/geoip2/) looks up a country or city in a MaxMind or DB-IP database file you download. + +geojson has [a class for every object in the GeoJSON spec](https://github.com/jazzband/geojson), and `geojson.dumps()` and `geojson.loads()` [wrap the standard json functions](https://github.com/jazzband/geojson#geojson-encodingdecoding) to encode and decode them. Check an object with [its `is_valid` property and `errors()` method](https://github.com/jazzband/geojson#validation). To encode your own classes the same way, [give them a `__geo_interface__`](https://github.com/jazzband/geojson#custom-classes), which GeoPandas' `from_features()` [also accepts](https://geopandas.org/en/stable/docs/reference/api/geopandas.GeoDataFrame.from_features.html). + +Latitude and longitude are angles, so measure distance and area in a projected coordinate system, in meters or feet. In GeoPandas, you [always need one](https://geopandas.org/en/stable/getting_started/introduction.html#Projections) for those operations, so reproject with `to_crs()` first. GeoDjango calls choosing a geometry field's SRID [an important decision](https://docs.djangoproject.com/en/stable/ref/contrib/gis/model-api/#selecting-an-srid), since projected systems ease distance calculations. geopy's `distance` works on an ellipsoidal model of the earth, so it takes latitudes and longitudes as they are. diff --git a/website/data/category_intros/gui-development.md b/website/data/category_intros/gui-development.md new file mode 100644 index 00000000..425ac8a8 --- /dev/null +++ b/website/data/category_intros/gui-development.md @@ -0,0 +1,33 @@ +Among Python GUI libraries, PySide6 is the default for desktop apps and tkinter for small tools. For a UI in the browser, pick NiceGUI. + +How to choose: + +- Full desktop app: PySide6, or PyQt6 if your app can be GPL or you buy a license +- Small tool without third-party packages: tkinter +- Modern look for a tkinter app: CustomTkinter +- tkinter layout drawn in Figma: Tkinter Designer +- Native widgets on Windows, macOS, and Linux: wxPython +- GNOME app on Linux: PyGObject +- GPU-rendered tools for your scripts: Dear PyGui +- Multi-touch apps on Android and iOS: Kivy +- Native widgets on desktop and mobile: Toga +- One codebase for web, desktop, and mobile: Flet +- Dashboards and web UIs: NiceGUI +- HTML/JavaScript frontend in a desktop window: pywebview +- GUI for an existing argparse script: Gooey + +PySide6 is Qt for Python, the [official Python bindings for Qt](https://doc.qt.io/qtforpython-6/), under the LGPL, the GPL, or a commercial license. Qt's docs [recommend a virtual environment](https://doc.qt.io/qtforpython-6/gettingstarted.html) over installing it into your system Python. Ship it with [pyside6-deploy](https://doc.qt.io/qtforpython-6/deployment/index.html). + +PyQt6 wraps the same Qt, but Riverbank [licenses it under the GPL or a commercial license, not the LGPL](https://www.riverbankcomputing.com/software/pyqt/). So a closed-source app needs a commercial PyQt6 license, while PySide6 can stay on the LGPL. + +tkinter is the standard Python interface to Tcl/Tk. Python's docs recommend the [themed tkinter.ttk widgets](https://docs.python.org/3/library/tkinter.ttk.html), which follow the platform's native theme, over the classic ones most online docs still use. + +NiceGUI runs a web server and shows your UI in the browser, which suits dashboards, micro web apps, and robotics projects. Pass `native=True` to `ui.run()` to [open it in a desktop window](https://nicegui.io/documentation/section_configuration_deployment) instead, or bundle it into an executable with nicegui-pack. + +Kivy runs the same code on Android, iOS, Linux, macOS, and Windows. Declare the widget tree in the [KV language](https://kivy.org/doc/stable/guide/lang.html) to keep the UI apart from your logic, and build Android packages with [Buildozer](https://kivy.org/doc/stable/guide/packaging-android.html). + +Toga [uses native system widgets, not themes](https://toga.beeware.org/en/stable/about/philosophy/), so a Toga app is a native app on each platform. Start with the [BeeWare tutorial](https://tutorial.beeware.org/), which packages your app with Briefcase. + +Flet builds web, desktop, and mobile apps from one Python codebase, [without HTML, CSS, or JavaScript](https://flet.dev/docs/). Package it for each platform with [flet build](https://flet.dev/docs/publish/). + +Whatever you pick, keep slow work out of event handlers, or the window freezes. The [tkinter docs](https://docs.python.org/3/library/tkinter.html) say to break it into smaller pieces with timers or run it in another thread, and Qt's docs suggest threads for the same reason. diff --git a/website/data/category_intros/hardware.md b/website/data/category_intros/hardware.md new file mode 100644 index 00000000..cd5adfad --- /dev/null +++ b/website/data/category_intros/hardware.md @@ -0,0 +1,13 @@ +Not every Python hardware library needs a board: pynput drives your keyboard and mouse, Bleak your Bluetooth LE devices. Jumpstarter automates hardware tests. + +How to choose: + +- Controlling or monitoring the keyboard and mouse: pynput +- Bluetooth Low Energy devices, like sensors: Bleak +- Automated tests on real or virtual hardware: Jumpstarter + +pynput [controls and monitors input devices](https://pynput.readthedocs.io/en/latest/): the mouse and the keyboard. To send input, create a `Controller` and [call its methods](https://pynput.readthedocs.io/en/latest/keyboard.html#controlling-the-keyboard), like `press()`, `release()`, or `type()` for a whole string. To react to input, open a `Listener` with your callbacks in a `with` block and [call `join()`](https://pynput.readthedocs.io/en/latest/keyboard.html#monitoring-the-keyboard). In a GUI app with its own main loop, call `start()` instead, so your code keeps running. + +Bleak is a [GATT client](https://bleak.readthedocs.io/en/latest/): it connects to Bluetooth Low Energy devices, like sensors, through one asynchronous, cross-platform API. Connect in an `async with BleakClient(...)` block and start your program with `asyncio.run()`. That's [the recommended way](https://bleak.readthedocs.io/en/latest/api/client.html#connecting-and-disconnecting), and the device disconnects even when your program is interrupted or raises. Scan the same way, in an [`async with BleakScanner(...)` block](https://bleak.readthedocs.io/en/latest/api/scanner.html#starting-and-stopping). + +Jumpstarter is an [open source framework for hardware-in-the-loop testing](https://jumpstarter.dev/main/introduction/index.html#introduction) on physical hardware and virtual devices. A person in `jmp shell`, a pytest script, and a CI pipeline all use the same APIs. [Local mode](https://jumpstarter.dev/main/introduction/index.html#local-mode) needs no Kubernetes and suits one developer with the hardware at hand. [Distributed mode](https://jumpstarter.dev/main/introduction/index.html#distributed-mode) runs a Kubernetes-based controller that leases devices, so teams can share them, including from CI. Write tests on the `JumpstarterTest` base class from jumpstarter-testing, which [handles the connection](https://jumpstarter.dev/main/getting-started/guides/examples/testing.html#the-jumpstartertest-base-class) for you. Set a `selector` for the device you need: the class connects from inside `jmp shell`, or leases a matching device outside it. diff --git a/website/data/category_intros/html-manipulation.md b/website/data/category_intros/html-manipulation.md new file mode 100644 index 00000000..84819777 --- /dev/null +++ b/website/data/category_intros/html-manipulation.md @@ -0,0 +1,21 @@ +Scraped HTML, however broken, parses in Beautiful Soup on lxml's parser. JustHTML packs a sanitizer into a Python HTML library and runs it by default. + +How to choose: + +- Pulling data out of scraped pages: Beautiful Soup, on lxml's parser +- Speed, XPath, or XML documents: lxml +- Sanitizing HTML your users submit: JustHTML +- XML you'd rather handle like JSON: xmltodict +- Escaping text you put into HTML: MarkupSafe + +Beautiful Soup is for [pulling data out of HTML and XML files](https://www.crummy.com/software/BeautifulSoup/bs4/doc/). It sits on a parser you pick and gives you one way to navigate, search, and change the tree. Its docs [recommend lxml as that parser](https://www.crummy.com/software/BeautifulSoup/bs4/doc/#installing-a-parser) for speed. Name the parser in the constructor, as in `BeautifulSoup(markup, "lxml")`. [Different parsers build different trees](https://www.crummy.com/software/BeautifulSoup/bs4/doc/#differences-between-parsers) from the same broken page, and your script should parse it the same way on every machine. The docs also say when to skip Beautiful Soup: [work directly atop lxml](https://www.crummy.com/software/BeautifulSoup/bs4/doc/#improving-performance) when computer time costs more than programmer time, and [parse with lxml](https://www.crummy.com/software/BeautifulSoup/bs4/doc/#css-selectors-through-the-css-property) when CSS selectors are all you need. + +lxml is [a Pythonic binding for libxml2 and libxslt](https://lxml.de/): the speed and XML features of those C libraries, with an API mostly compatible with ElementTree's. For HTML, use [lxml.html](https://lxml.de/lxmlhtml.html), which adds HTML-specific methods to lxml's elements, and every element takes `.xpath()`. On broken pages, lxml's docs say you often need only [Beautiful Soup's encoding detection](https://lxml.de/lxmlhtml.html#really-broken-pages). Leave the rest to lxml's own parser, which is several times faster. XPath has the same injection problem as SQL. When a value comes from outside, [pass it as an XPath variable](https://lxml.de/FAQ.html#how-do-i-use-lxml-safely-as-a-web-service-endpoint) instead of formatting it into the expression. + +JustHTML [parses HTML like a browser](https://github.com/EmilStenstrom/justhtml), including broken markup. `JustHTML(html)` [sanitizes by default](https://emilstenstrom.github.io/justhtml/sanitization.html) against a strict allowlist. The output is safe in a page body, but [not automatically safe inside a ` -
-
-
-

Search every project in one place

-
-

- Press / to search. Tap a tag to filter. Click any row for - details. -

-
- -
-

Search and filter

-
- - - - - -
-
- Filtering for - -
-
- -

Results

-
- - - - - - - - - - - - - - {% for entry in entries %} +{% macro entry_rows(entry, index) %} + {# On section and subcategory pages the group heading already names the row's use case #} + {% set grouped_page = entry_groups and not group_categories %} + {% set section_path = "/categories/" ~ parent_category.slug ~ "/" if parent_category else current_path %} - + {% if entry.description %} - + {% endif %} - + +{% endmacro %} +
+
+ {% if entry_groups %} + {% set named_groups = entry_groups | selectattr("name") | list %} +

+ {% if group_categories %}Listed by section, in editorial order.{% elif named_groups %}Listed in editorial order, grouped by use case.{% else %}Listed in editorial order.{% endif %} + Click a column to re-sort the whole list. +

+ {% else %} +
+

{{ entries | length }} {{ category.name }} projects

+
+ {% endif %} +

+ Press / to search. Tap a tag to filter. Click any row for + details. +

+
+ +
+

Search and filter

+
+ + + + + +
+
+ +

Results

+
+
Row number - - - - - - - - Tags - -
{{ entry.description | safe }}
@@ -270,7 +206,107 @@
+ + + + + + + + + + + + + {% if entry_groups %} + {% set row = namespace(index=0) %} + {% for group in entry_groups %} + {% if group.name %} + + + + {% endif %} + {% for entry in group.entries %} + {% set row.index = row.index + 1 %} + {{ entry_rows(entry, row.index) }} {% endfor %} + {% endfor %} + {% else %} + {% for entry in entries %} + {{ entry_rows(entry, loop.index) }} + {% endfor %} + {% endif %}
Row number + + + + + + + + Tags + +
+

+ {% if group.url %}{{ group.name }}{% else %}{{ group.name }}{% endif %} + {{ group.entries | length }} project{{ "s" if group.entries | length != 1 }} +

+
@@ -284,6 +320,15 @@
+{% if guide_html %} +
+
+

{{ category.name }} guide

+
{{ guide_html | safe }}
+
+
+{% endif %} +
diff --git a/website/templates/index.html b/website/templates/index.html index b5540624..e4d95d10 100644 --- a/website/templates/index.html +++ b/website/templates/index.html @@ -96,7 +96,6 @@
{% endif %} -
@@ -132,12 +131,6 @@ aria-label="Search projects" />
-
- Filtering for - -

Results

@@ -175,7 +168,6 @@ {% for entry in entries %} {% for subcat in entry.subcategories %} - + {{ subcat.name }} {% endfor %} {% for cat in entry.categories %} {{ cat }} {% endfor %} {{ entry.groups[0] }} @@ -245,8 +233,6 @@ Stdlib diff --git a/website/templates/redirect.html b/website/templates/redirect.html new file mode 100644 index 00000000..292c443c --- /dev/null +++ b/website/templates/redirect.html @@ -0,0 +1,11 @@ + + + +Redirecting… + + + + +

Redirecting…

+Click here if you are not redirected. + diff --git a/website/tests/test_build.py b/website/tests/test_build.py index 2f0af663..3588a719 100644 --- a/website/tests/test_build.py +++ b/website/tests/test_build.py @@ -17,6 +17,7 @@ from build import ( detect_source_type, extract_entries, extract_github_repo, + load_category_intro, load_downloads, load_pypi_badges, load_stars, @@ -279,8 +280,7 @@ class TestBuild: parser.feed(category_html) assert 'href="/categories/widgets/"' in index_html - assert 'data-value="Widgets"' in index_html - assert parser.title.strip() == "Widgets Python Libraries - Awesome Python" + assert parser.title.strip() == "Python Widgets Libraries - Awesome Python" assert parser.meta_by_name["description"] == "Widget libraries. Also see awesome-widgets. Explore 2 curated Python projects in Widgets." assert parser.links_by_rel["canonical"] == "https://awesome-python.com/categories/widgets/" assert parser.meta_by_property["og:url"] == "https://awesome-python.com/categories/widgets/" @@ -291,10 +291,45 @@ class TestBuild: assert 'href="https://example.com/w1"' in category_html assert "A widget." in category_html assert 'href="https://github.com/owner/w2"' in category_html - assert '' in category_html + assert '
' in category_html assert "42" in category_html assert "2026-01-01T00:00:00+00:00" in category_html + def test_build_links_description_anchors_to_category_pages(self, tmp_path): + readme = textwrap.dedent("""\ + # Awesome Python + + Intro. + + ## Projects + + **Tools** + + ## Audio & Video + + _Media tools._ + + - [a1](https://example.com/a1) - A media tool. + + ## Widgets + + _Widget libraries. Also see [Audio & Video](#audio--video) and [Gadgets](#gadgets)._ + + - [w1](https://example.com/w1) - A widget. + + # Contributing + + Help! + """) + (tmp_path / "README.md").write_text(readme, encoding="utf-8") + self._copy_real_templates(tmp_path) + + build(tmp_path) + + category_html = (tmp_path / "website" / "output" / "categories" / "widgets" / "index.html").read_text(encoding="utf-8") + assert 'Also see Audio & Video' in category_html + assert 'Gadgets' in category_html + def test_build_creates_llms_text_alternate_without_sponsors(self, tmp_path): readme = textwrap.dedent("""\ # Awesome Python @@ -622,7 +657,7 @@ class TestBuild: assert set(graph) == {"WebSite", "CollectionPage", "BreadcrumbList"} assert graph["WebSite"]["@id"] == "https://awesome-python.com/#website" collection = graph["CollectionPage"] - assert collection["name"] == "Widgets Python Libraries" + assert collection["name"] == "Python Widgets Libraries" assert collection["@id"] == "https://awesome-python.com/categories/widgets/" assert collection["url"] == "https://awesome-python.com/categories/widgets/" assert collection["description"] == "Widget libraries. Explore 2 curated Python projects in Widgets." @@ -672,7 +707,7 @@ class TestBuild: graph = {node["@type"]: node for node in data["@graph"]} collection = graph["CollectionPage"] - assert collection["name"] == "AI & ML Python Libraries" + assert collection["name"] == "Python AI & ML Libraries" assert collection["@id"] == "https://awesome-python.com/categories/ai-ml/" assert collection["url"] == "https://awesome-python.com/categories/ai-ml/" assert collection["description"] == "Explore 1 curated Python projects in AI & ML. Part of the Awesome Python catalog." @@ -815,75 +850,6 @@ class TestBuild: {"@type": "ListItem", "position": 2, "name": "Sponsorship", "item": "https://awesome-python.com/sponsorship/"}, ] - def test_index_embeds_filter_urls_json(self, tmp_path): - readme = textwrap.dedent("""\ - # T - - ## Projects - - **AI & ML** - - ## Deep Learning - - - [dl1](https://example.com/dl1) - DL. - - ## Machine Learning - - - Classical - - - [ml1](https://example.com/ml1) - ML. - - # Contributing - - Done. - """) - self._copy_real_templates(tmp_path) - (tmp_path / "README.md").write_text(readme, encoding="utf-8") - build(tmp_path) - - site = tmp_path / "website" / "output" - index_html = (site / "index.html").read_text(encoding="utf-8") - - marker = '", start) - data = json.loads(index_html[start:end]) - - assert data["Deep Learning"] == "/categories/deep-learning/" - assert data["Machine Learning"] == "/categories/machine-learning/" - assert data["AI & ML"] == "/categories/ai-ml/" - assert data["Machine Learning > Classical"] == "/categories/machine-learning/classical/" - - def test_filter_urls_json_escapes_closing_script_tag(self, tmp_path): - readme = textwrap.dedent("""\ - # T - - ## Projects - - ## Sneaky - - - [a](https://example.com) - A. - - # Contributing - - Done. - """) - self._copy_real_templates(tmp_path) - (tmp_path / "README.md").write_text(readme, encoding="utf-8") - build(tmp_path) - - site = tmp_path / "website" / "output" - index_html = (site / "index.html").read_text(encoding="utf-8") - - marker = '", start) - block = index_html[start:end] - assert "" not in block - data = json.loads(block) - assert any("Sneaky" in key for key in data) - def test_build_creates_group_pages(self, tmp_path): readme = textwrap.dedent("""\ # T @@ -924,7 +890,7 @@ class TestBuild: assert "wf1" in web_dev assert "dl1" not in web_dev - def test_tag_buttons_have_data_url(self, tmp_path): + def test_tags_link_to_pages_and_subcategory_anchors(self, tmp_path): readme = textwrap.dedent("""\ # T @@ -949,11 +915,176 @@ class TestBuild: site = tmp_path / "website" / "output" index_html = (site / "index.html").read_text(encoding="utf-8") - assert 'data-value="Deep Learning"' in index_html - assert 'data-url="/categories/deep-learning/"' in index_html - assert 'data-value="AI & ML"' in index_html or 'data-value="AI & ML"' in index_html - assert 'data-url="/categories/ai-ml/"' in index_html - assert 'data-url="/categories/deep-learning/vision/"' in index_html + category_html = (site / "categories" / "deep-learning" / "index.html").read_text(encoding="utf-8") + + for html in (index_html, category_html): + assert 'href="/categories/deep-learning/#vision"' in html + assert 'href="/categories/ai-ml/"' in html + assert "data-url=" not in html + assert 'href="/categories/deep-learning/"' in index_html + assert 'id="vision"' in category_html + + _REDIRECT_README = textwrap.dedent("""\ + # Awesome Python + + Intro. + + ## Projects + + **Tools** + + ### Widgets + + - [w1](https://example.com/w1) - A widget. + + ## Contributing + + Help! + """) + + def _write_redirects(self, tmp_path, redirects): + data_dir = tmp_path / "website" / "data" + data_dir.mkdir(parents=True) + (data_dir / "redirects.json").write_text(json.dumps(redirects), encoding="utf-8") + + def test_build_writes_redirect_stub_outside_sitemap(self, tmp_path): + self._copy_real_templates(tmp_path) + (tmp_path / "README.md").write_text(self._REDIRECT_README, encoding="utf-8") + self._write_redirects(tmp_path, {"/categories/old-widgets/": "/categories/widgets/"}) + build(tmp_path) + + site = tmp_path / "website" / "output" + stub = (site / "categories" / "old-widgets" / "index.html").read_text(encoding="utf-8") + assert '' in stub + assert '' in stub + assert '' in stub + assert "old-widgets" not in (site / "sitemap.xml").read_text(encoding="utf-8") + + def test_build_renders_category_intro_and_uses_lead_as_meta_description(self, tmp_path): + self._copy_real_templates(tmp_path) + (tmp_path / "README.md").write_text(self._REDIRECT_README, encoding="utf-8") + intros_dir = tmp_path / "website" / "data" / "category_intros" + intros_dir.mkdir(parents=True) + (intros_dir / "widgets.md").write_text("Use `w1` for most apps.\n\nSee [the docs](https://example.com/docs).\n\nHow to choose:\n\n- Small apps: w1\n", encoding="utf-8") + build(tmp_path) + + category_html = (tmp_path / "website" / "output" / "categories" / "widgets" / "index.html").read_text(encoding="utf-8") + parser = HeadMetadataParser() + parser.feed(category_html) + assert parser.meta_by_name["description"] == "Use w1 for most apps." + assert '

Use w1 for most apps.

' in category_html + assert "
  • Small apps: w1
  • " in category_html + assert 'the docs' in category_html + + def test_build_renders_category_guide_below_table(self, tmp_path): + self._copy_real_templates(tmp_path) + (tmp_path / "README.md").write_text(self._REDIRECT_README, encoding="utf-8") + intros_dir = tmp_path / "website" / "data" / "category_intros" + intros_dir.mkdir(parents=True) + (intros_dir / "widgets.md").write_text("Use w1.\n\nHow to choose:\n\n- Small apps: w1\n\nSet up w1 once per process.\n", encoding="utf-8") + build(tmp_path) + + category_html = (tmp_path / "website" / "output" / "categories" / "widgets" / "index.html").read_text(encoding="utf-8") + intro_html = category_html.split('
    ', 1)[1].split("
    ", 1)[0] + assert "
  • Small apps: w1
  • " in intro_html + assert "Set up w1" not in intro_html + guide_html = category_html.split('
    ', 1)[1] + assert "

    Widgets guide

    " in guide_html + assert "

    Set up w1 once per process.

    " in guide_html + assert category_html.index('id="guide"') > category_html.index("
    ") + assert 'Widgets guide' in category_html + + def test_section_page_groups_rows_by_use_case_in_readme_order(self, tmp_path): + readme = textwrap.dedent("""\ + # T + + ## Projects + + **Tools** + + ### Widgets + + - Small + - [w2](https://example.com/w2) - Second. + - [w1](https://example.com/w1) - First. + - Large + - [w3](https://example.com/w3) - Third. + - [sqlite3](https://docs.python.org/3/library/sqlite3.html) - Stdlib. + + # Contributing + + Done. + """) + self._copy_real_templates(tmp_path) + (tmp_path / "README.md").write_text(readme, encoding="utf-8") + build(tmp_path) + + site = tmp_path / "website" / "output" / "categories" + html = (site / "widgets" / "index.html").read_text(encoding="utf-8") + assert 'data-default-sort="editorial"' in html + positions = [html.index(marker) for marker in ('

    ', ">w2w1', ">w3Small' in html + assert 'Small' in html + assert '' in html + assert '' in html + + subcategory_html = (site / "widgets" / "small" / "index.html").read_text(encoding="utf-8") + assert 'data-default-sort="editorial"' in subcategory_html + assert "group-row" not in subcategory_html + assert subcategory_html.index(">w2w1Machine Learning') + dl_heading = html.index('Deep Learning') + assert ml_heading < html.index(">ml1dl1ml1Use w1 for most apps.

    \n

    How to choose:

    \n
      \n
    • Small apps: w1
    • \n
    • Big apps: w2
    • \n
    \n" + assert guide_html == "

    Configure w1 once.

    \n

    Pin w2.

    \n" + assert lead == "Use w1 for most apps." + + def test_keeps_everything_above_table_without_how_to_choose_list(self, tmp_path): + path = tmp_path / "widgets.md" + path.write_text("Use w1.\n\n- Small apps: w1\n\nConfigure w1 once.\n", encoding="utf-8") + intro_html, guide_html, _ = load_category_intro(path) + assert "Configure w1 once." in intro_html + assert guide_html == "" + + def test_returns_empty_strings_without_intro_file(self, tmp_path): + assert load_category_intro(tmp_path / "missing.md") == ("", "", "")