Commit Graph
288 Commits
Author SHA1 Message Date
Vinta ChenandClaude 62a4e7c4e5 fix: make tags link to category pages instead of filtering in place
Clicking a category or group tag filtered the homepage table in place while rewriting the address bar to the category URL, so one URL rendered two different pages and readers never reached category pages with their intros and guides; category pages also showed a redundant self-referential filter bar.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:34:07 +08:00
Vinta ChenandClaude e5088a6f5c docs: rewrite Data Validation category intro
Covers Pydantic for API input and config, Pandera for dataframes, and jsonschema for JSON Schema validation.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:27:34 +08:00
Vinta ChenandClaude c52f21f99d docs: rewrite CLI Development category intro
Covers argparse for basic apps, Click or Typer beyond that, Rich for output, and TUI frameworks.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:27:29 +08:00
Vinta ChenandClaude ee2874a463 docs: rewrite Testing category intro
Covers pytest as the default, a how-to-choose item per README subcategory, and a guide on Hypothesis, Playwright, tox/Nox, mocks, and coverage.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:27:25 +08:00
Vinta ChenandClaude 3c2afa13bb docs: rewrite Audio & Video Processing category intro
Covers librosa for audio analysis, MoviePy for scripted video editing, VidGear for real-time video, Mutagen vs tinytag for tag read/write and licensing, and beets as a CLI tag/organization tool.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:27:20 +08:00
Vinta ChenandClaude 88cfcf6f4d docs: rewrite Computer Vision category intro
Covers OpenCV as the default, Ultralytics YOLO for detection/segmentation/pose models and its AGPL-3.0/Enterprise License terms, Kornia for GPU-batch vision ops, FiftyOne for dataset curation, and pytesseract vs EasyOCR for OCR.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:27:02 +08:00
Vinta ChenandClaude a9bb827b4a docs: add Code Analysis category intro
Ruff for linting and formatting, a type checker alongside it, and pre-commit to run them, with per-tool guidance sourced from each project's own docs.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:17:54 +08:00
Vinta ChenandClaude 5216fb3c1a docs: add Image Processing category intro
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:17:34 +08:00
Vinta ChenandClaude 2930319d48 docs: add CMS category intro
Explains when to pick Wagtail (developer-defined page types) versus
django CMS (editors composing pages live), based on each project's
own documentation.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-27 00:16:51 +08:00
Vinta ChenandClaude 2b9469c0ae feat: group category page rows by use case, restructure intro
Category pages sorted rows by downloads, which buried editorial leads (the Django ORM showed as row 7 of 7, tkinter as 14 of 14) against CONTRIBUTING.md's "position is the marker"; the full intro in the hero pushed the table 2.5 screens down on a phone; and every page repeated the same "Search every project in one place" H2.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 23:33:57 +08:00
Vinta ChenandClaude 528f9edbb7 fix: expose entry description rows to screen readers
Every entry description row carried aria-hidden="true", hiding descriptions that sighted readers can see from screen readers.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 23:32:20 +08:00
Vinta ChenandClaude 9cc721f533 fix: match category meta titles to search query word order
Category page titles read "ORM Python Libraries", but people search "python orm", so the title word order didn't match the query.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 23:28:51 +08:00
Vinta ChenandClaude eb51124608 refactor: extract category table row markup into entry_rows macro
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 22:20:37 +08:00
Vinta ChenandClaude eaddf9441e docs: drop links from gui-development how-to-choose list
7 of its 13 items carried links while the other intros' lists carry none, and the entry table already links every project.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 22:14:39 +08:00
Vinta ChenandClaude 00915b0300 docs: trim ai-and-agents how-to-choose list to 12 items
The list had 22 items, which in the upcoming layout sits above the entry table and would push it 4-5 phone screens down; it now has one item per README subcategory, in README order, reusing the intro's own wording.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 22:14:15 +08:00
Vinta ChenandClaude 2cc1bb440d docs: rewrite ORM category intro
The old intro told every FastAPI app to use SQLModel, which SQLModel's own docs don't claim beyond simple cases, and leaned on API names from SQLAlchemy's 2.0 rename (Mapped, mapped_column).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 22:09:22 +08:00
Vinta ChenandClaude 4d2c9906e3 docs: rewrite GUI Development category intro
The old intro's lead used the same "For a Python X library, use …" template as other category pages, and it carried version-bound usage tips (Qt Widgets vs Qt Quick, pyside6-uic, per-toolkit threading helpers) instead of each project's recommended setup.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 22:09:16 +08:00
Vinta ChenandClaude 17bb6270d6 docs: add AI and Agents category intro
The AI and Agents category page had no intro, so readers got a 35-project list with no guidance on which library to pick for building an agent, serving a model, or fine-tuning.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 22:09:10 +08:00
Vinta ChenandClaude 1dfa95834d docs: add GUI Development category intro
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 20:41:18 +08:00
Vinta ChenandClaude 3870b0a723 docs: add Audio & Video Processing category intro
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 20:40:53 +08:00
Vinta ChenandClaude b71151a808 docs: add Computer Vision category intro
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 20:40:29 +08:00
Vinta ChenandClaude ec54348e4c docs: add CLI Development category intro
Covers recommended usage from each project's own docs for Typer, argparse, Click, Rich, and Textual.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 20:10:52 +08:00
Vinta ChenandClaude 7d19f84e63 docs: add Data Validation category intro
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 20:10:14 +08:00
Vinta ChenandClaude fd300c4453 docs: add Testing category intro
Recommends pytest, Hypothesis, Playwright, and tox/nox based on each project's own docs.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 20:09:42 +08:00
Vinta ChenandClaude a74ca52860 feat: render per-category intro text above the entry list
Category pages carried no text of their own beyond the README one-line description, and most meta descriptions fell back to a generic "Explore N curated Python projects" line, which correlated with weak search rankings for category queries. This adds optional per-category intro markdown files rendered under the H1, with the first paragraph used as the meta description and links opening in a new tab, starting with the ORM category.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 19:42:22 +08:00
Vinta ChenandClaude c10264977e feat: serve redirect stubs for renamed and dissolved category slugs
Renamed or dissolved category slugs (e.g. /categories/web-servers/rpc/, /categories/code-analysis/code-linters/) returned 404, Search Console lists 7 of them, and each audit re-home was dropping the old URL's ranking.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-26 18:45:53 +08:00
Vinta ChenandClaude b7313dcffc test: assert bundled entries have a null PyPI override
The uv-audit bug had no automated guard: pypi_name_overrides.json is
a manual registry, so a wrong-package mapping is only caught if
someone already suspects it.

Two broader checks were measured against the real list and rejected.
Checking that PyPI metadata links back to the entry's GitHub repo
would not have caught uv-audit, since that package declares no
home_page or project_urls, landing it in a 26-entry bucket of
packages that simply don't declare a repo (numba, selenium, pyglet,
etc.), plus 10 benign cases of orgs moving or splitting bindings.
Flagging display-name/repo-name mismatches yields 46 hits, all
legitimate python-X-repo-to-X-package pairs, with uv-build sitting
among them despite being a real Astral package with the identical
shape to uv-audit.

What discriminates is the bundled marker itself: a "(part of X)"
entry ships inside something else and has no package of its own, so
the sweep must never query it. This test walks the real README and
requires a null override for every bundled entry whose normalized
name is PyPI-shaped. Verified it fails with exactly the uv-audit
message when that override is removed, and passes with it restored,
across the three current bundled entries with no false positives. It
runs offline, fitting the existing network-less CI.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-23 01:46:13 +08:00
Vinta ChenandClaude 8e7d2bc62c fix: add null pypi override for uv-audit
Renaming the entry from "uv audit" to "uv-audit" made the name PyPI-shaped: normalize() leaves spaces alone, so "uv audit" failed PYPI_NAME_RE and collect_names skipped it, but "uv-audit" passes, so the next sweep would have queried PyPI for it.

A uv-audit package does exist on PyPI, but it is version 0.1.9 by Alekse Marusich of rocshers, an unrelated third-party tool whose summary ("uv Tool for checking dependencies for vulnerabilities") is close enough to be mistaken for Astral's built-in uv audit subcommand. Without the override the entry would have shown that stranger's download count and lost its Bundled badge.

The sweep now writes uv-audit as NOT_FOUND, which load_downloads skips, so the badge is unaffected.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-23 01:30:53 +08:00
Vinta ChenandClaude 0f677911f0 feat: add "Multiple on PyPI" badge for SDK monorepos
azure-sdk-for-python and google-cloud-python were rendering "Not on
PyPI", which is misleading. Both do ship on PyPI, just as many
per-service packages (azure-identity, azure-storage-blob,
google-cloud-storage, etc.) rather than under the repo name.

pypi_name_overrides.json already recorded that distinction in its
reason field; those two entries now carry an optional "badge" value
that build.py reads into the PyPI Downloads column. The other sixteen
no-count entries (cpython, renpy, agent skill repos, etc.) keep
"Not on PyPI" since that remains accurate for them.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-23 01:28:33 +08:00
Vinta ChenandClaude 733021ae5b feat: add Bundled badge for entries shipped inside a larger project
django.db.models, geodjango, httpx.URL and uv audit were rendering
"Not on PyPI" alongside eighteen genuinely standalone projects that
simply are not packaged on PyPI, conflating two different reasons for
a missing download count.

These four entries now carry a "(part of X)" description prefix in
README.md, mirroring the existing "(Python standard library)"
convention. build.py reads that prefix into a bundled flag that both
templates render as a "Bundled" badge.

The prefix approach was chosen over a separate data file so README.md
stays the single source of content truth, and over inferring from the
entry name because geodjango is neither dotted nor spaced and would
have been missed. Redundant tail wording was trimmed from the
httpx.URL, geodjango and uv audit descriptions now that the prefix
names the parent.

The new entry format is documented in CONTRIBUTING.md and the
vocabulary in CONTEXT.md.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-23 01:09:20 +08:00
Vinta ChenandClaude 2fc0edd565 feat: rename Built-in source type to Stdlib and show it in downloads column
Standard-library entries rendered "Not on PyPI" in the PyPI Downloads
column, which read like missing data rather than a deliberate category
— the build already forces downloads to None for them so they never
pick up a same-named PyPI backport. They now render a "Stdlib" badge
instead, while genuine non-PyPI entries keep "Not on PyPI".

The filter tag is renamed to match, so the source-type value, the row
filter tag, and the synthetic category heading all read Stdlib now.
The literal "Built-in" strings scattered through build.py are routed
through the existing BUILTIN_FILTER constant so the label lives in one
place. The category page slug stays "built-in" so the public URL
/categories/built-in/ does not break.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-23 01:01:17 +08:00
Vinta ChenandClaude 4fc35351b2 docs: update og-image wordings to match current site copy
Kicker now mirrors the hero kicker ("The definitive list that
answers..."), replacing the old "field guide" line changed in
8b14b4f. Subtitle now matches the current tagline ("An opinionated
guide to the best Python frameworks, libraries, and tools.") on one
line. PNG regenerated from the SVG.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 20:31:39 +08:00
Vinta Chen 8b14b4fd1e update styles 2026-08-16 20:25:19 +08:00
Vinta Chen ee8089d0d5 add build date to llms.txt 2026-08-16 20:16:05 +08:00
Vinta ChenandClaude 30c9f18bc2 feat: annotate llms.txt entries with PyPI download counts
Entries now end with "(PyPI downloads/month: N, GitHub stars: M)" where
known, replacing the stars-only note, since download counts are the
list's stated primary evidence signal. annotate_entries_with_stars is
renamed to annotate_entries_with_stats and looks downloads up by the
first link's display name, skipping category-index bullets (which link
into the site itself) and Built-in entries (which would otherwise hit
same-named PyPI backports like logging or asyncio).

The intro now mirrors the README subtitle verbatim with the
project/category totals on their own line below, and the "opinionated
catalog" wording is gone since the shortlist ADR is literally titled
"shortlist, not a catalog".

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:43:52 +08:00
Vinta ChenandClaude bfdd087282 feat: show "Not on PyPI" badge for missing download counts
Replaces the em dash in the PyPI Downloads column with a source-badge
pill labeled "Not on PyPI", reusing the existing badge style used by
the stars column for visual consistency. Sorting is unaffected since
non-numeric cells already parse as missing.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:20:55 +08:00
Vinta ChenandClaude f5f7a1c50c refactor: swap PyPI Downloads and GitHub Stars column order
Downloads is now the default sort, so it sits directly after the
project name in both the index and category table templates. The
source-type badge stays in the stars cell.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:19:42 +08:00
Vinta ChenandClaude da491ff8cc docs: rename Downloads/Month column header to PyPI Downloads
The header no longer carries the per-month unit; expand-row text keeps
its 'downloads/month' wording.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:16:03 +08:00
Vinta ChenandClaude d3fec1f7fa feat: default-sort the website table by downloads per month
Entries with a download count now sort first (descending), with
stars, then Built-in, then name as fallback tiers for entries that
lack a count. main.js mirrors this in its default activeSort, clean
URL check, and third-click reset target. Sorting by stars remains one
header click away.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:15:21 +08:00
Vinta ChenandClaude b841db68ba refactor: make pypi_name_overrides.json entries self-documenting
Every entry is now {"package": str|null, "reason": str|null} instead of
a bare string/null. Reasons are required for null packages, explaining
why the name must never be queried (squatted name, stdlib module,
monorepo umbrella, GitHub-only project, and so on). Reasons are
optional for remaps and kept only on the six non-obvious ones: pytorch
(squatter), jinja (jinja is Jinja1), strawberry (unrelated bookmarking
service), django-rules (abandoned fork), django-rest-framework (dead
alias), and devpi (deprecated metapackage); plain publishes-as-X
remaps get a null reason.

load_overrides() in the clickpy fetcher now extracts the package field
from each entry; resolve() and the pepy/bigquery cross-check scripts
are unchanged since they consume load_overrides()'s output.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:08:55 +08:00
Vinta ChenandClaude fcfc65df00 data: record all remaining NOT_FOUND names as explicit null overrides
Every queried name now resolves 447/447. Adds 23 explicit null
overrides so squatters can never silently attach a PyPI number to
these names later: stdlib-named entries (concurrent-futures, difflib,
mimetypes, sqlite3, tkinter, tomllib, zoneinfo), interpreters
(micropython, pypy), monorepo umbrellas (azure-sdk-for-python,
google-cloud-python), self-hosted or distro-installed projects (odoo,
cloud-init, warehouse), GitHub-only projects (thealgorithms,
geodjango, django-db-models, django-ai-plugins, graphify,
sentry-skills, social-engineer-toolkit, trailofbits-skills), and
httpx-url (a class within httpx, not a package).

Caveat: graphify and django-ai-plugins are young projects that may
legitimately publish to PyPI later — flip their null to a remap
during a future audit if they do.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 18:57:13 +08:00
Vinta ChenandClaude f3ff733fd7 fix: add three more PyPI name remaps for NOT_FOUND triage
autobahn-python publishes as autobahn (7.1M/mo), pangu-py as pangu, and
strawberry-django as strawberry-graphql-django (1.5M/mo). httpx.URL is
left unmapped deliberately since it's a class within the httpx package,
not a package of its own.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 18:54:58 +08:00
Vinta ChenandClaude cc804b1de2 feat: add Downloads/Month column to website table
Sourced from website/data/pypi_downloads.tsv the same way
github_stars.json feeds the stars column. The new sortable column
sits between GitHub Stars and Last Commit on the homepage and
category pages, formatted with thousands separators like stars, with
an em dash when no PyPI data exists. Rows are matched by normalized
README display name; Built-in entries never show counts since
same-named PyPI packages are stdlib backports (e.g. the asyncio
package).

Below 960px the column hides and the count moves into the expand
row, mirroring the existing Last Commit treatment. main.js gains the
downloads sort branch and URL param.

The deploy workflow fetches the TSV via the new
make fetch_pypi_downloads target with a daily actions/cache
fallback, mirroring the stars fetch, but non-fatal: the column
degrades to dashes when the fetch fails, unlike stars which the
build requires.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 17:52:04 +08:00
Vinta ChenandClaude b440e64cd0 fix: resolve PyPI download counts through curated package overrides
A pypi.org identity sweep of all 438 cached rows (project_urls/home_page
vs entry GitHub URL) found download counts were looked up by README
display name, so entries whose name differs from the canonical package
silently measured squatters or dead predecessors: pytorch measured a
squatter (169,737/mo vs torch's 94M), jinja measured Jinja1 (3,168 vs
jinja2's 736M), django-rest-framework a dead alias package (real:
djangorestframework), django-rules an abandoned fork (real: rules),
strawberry an unrelated bookmarking service (real: strawberry-graphql),
devpi a deprecated metapackage (mapped to devpi-server).

New curated website/data/pypi_name_overrides.json maps normalized
README name to the real package, or null for projects not
pip-installable whose name is squatted or a relic (cpython, pyenv,
renpy, python-patterns, winpython); also maps mem0 to mem0ai, fasthtml
to python-fasthtml, and playwright-python to playwright.

All three fetch scripts resolve names through it; the clickpy TSV
cache gains a package column recording what each row actually
measured. .gitignore switches website/data/ to website/data/* with a
negation so the curated overrides file is tracked while caches stay
ignored.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 17:51:11 +08:00
Vinta ChenandClaude 2c2ce2fcac docs: document empirically verified README parsing behavior
Capture parser quirks worth knowing before editing README.md:
everything above  is ignored, new subcategories need no
parser change, a standalone all-bold paragraph becomes a Thematic
Group marker, prose after  leaks into llms.txt, and the
build's "Total entries" figure counts sub-items rather than just
entries.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:36 +08:00
Vinta ChenandClaude 2c50d9e63a docs: document gitignored output and orphan-key behavior
Note that data/github_stars.json is gitignored and fetched by CI at
deploy time, so local runs are preview-only and should never be
committed; entries removed from README.md just leave harmless orphan
keys behind.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:31 +08:00
Vinta ChenandClaude 39eeab7dff docs: warn about wrong-package PyPI download counts
The pypi downloads sweep looks up counts by README display name; when
the display name differs from the canonical package, the row silently
measures an unrelated squatter or a dead predecessor. Document the
failure mode in both the fetcher's docstring and the audit skill so
famous entries with off-looking counts get identity-verified before
being cited.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:15 +08:00
Vinta ChenandClaude a2da210345 feat: stamp pypi downloads cache with fetched_at and header row
Adds a header row (name, downloads, fetched_at) to data/pypi_downloads.tsv
and stamps every row with the sweep date, so audits can tell evidence age
and skip re-fetching when the cache is less than 7 days old. The sweep
itself costs ~1s, so freshness is checked by the reader (audit-the-list
skill) instead of skip logic in the fetch script. SKILL.md documents the
7-day freshness rule.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 02:58:20 +08:00
Vinta ChenandClaude 1abfb1af42 refactor: split download fetchers into writer and cross-check scripts
Authored by the parallel fetch-scripts session (its intended 0cffb5f
never landed; the content rode into an audit commit by accident and is
extracted here): fetch_pypi_downloads_via_clickpy.py becomes the
flagless full-README sweep and sole writer of data/pypi_downloads.tsv;
new fetch_pypi_downloads_via_bigquery.py is a print-only cross-check
taking explicit names behind a 400 GB billing cap; audit-the-list
SKILL.md step 2 updated to match.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:57:16 +08:00
Vinta ChenandClaude 5297061a4f refactor: rename download-fetch scripts to source-suffixed scheme
fetch_pypi_downloads.py becomes fetch_pypi_downloads_via_clickpy.py and
fetch_pepy_downloads.py becomes fetch_pypi_downloads_via_pepy.py, ahead
of splitting the BigQuery path into its own file. Updates the usage
strings, the cross-file import, and the audit-the-list skill's
references to match. Pure rename, no behavior change.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:35:24 +08:00