Commit Graph
251 Commits
Author SHA1 Message Date
Vinta ChenandClaude da491ff8cc docs: rename Downloads/Month column header to PyPI Downloads
The header no longer carries the per-month unit; expand-row text keeps
its 'downloads/month' wording.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:16:03 +08:00
Vinta ChenandClaude d3fec1f7fa feat: default-sort the website table by downloads per month
Entries with a download count now sort first (descending), with
stars, then Built-in, then name as fallback tiers for entries that
lack a count. main.js mirrors this in its default activeSort, clean
URL check, and third-click reset target. Sorting by stars remains one
header click away.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:15:21 +08:00
Vinta ChenandClaude b841db68ba refactor: make pypi_name_overrides.json entries self-documenting
Every entry is now {"package": str|null, "reason": str|null} instead of
a bare string/null. Reasons are required for null packages, explaining
why the name must never be queried (squatted name, stdlib module,
monorepo umbrella, GitHub-only project, and so on). Reasons are
optional for remaps and kept only on the six non-obvious ones: pytorch
(squatter), jinja (jinja is Jinja1), strawberry (unrelated bookmarking
service), django-rules (abandoned fork), django-rest-framework (dead
alias), and devpi (deprecated metapackage); plain publishes-as-X
remaps get a null reason.

load_overrides() in the clickpy fetcher now extracts the package field
from each entry; resolve() and the pepy/bigquery cross-check scripts
are unchanged since they consume load_overrides()'s output.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:08:55 +08:00
Vinta ChenandClaude fcfc65df00 data: record all remaining NOT_FOUND names as explicit null overrides
Every queried name now resolves 447/447. Adds 23 explicit null
overrides so squatters can never silently attach a PyPI number to
these names later: stdlib-named entries (concurrent-futures, difflib,
mimetypes, sqlite3, tkinter, tomllib, zoneinfo), interpreters
(micropython, pypy), monorepo umbrellas (azure-sdk-for-python,
google-cloud-python), self-hosted or distro-installed projects (odoo,
cloud-init, warehouse), GitHub-only projects (thealgorithms,
geodjango, django-db-models, django-ai-plugins, graphify,
sentry-skills, social-engineer-toolkit, trailofbits-skills), and
httpx-url (a class within httpx, not a package).

Caveat: graphify and django-ai-plugins are young projects that may
legitimately publish to PyPI later — flip their null to a remap
during a future audit if they do.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 18:57:13 +08:00
Vinta ChenandClaude f3ff733fd7 fix: add three more PyPI name remaps for NOT_FOUND triage
autobahn-python publishes as autobahn (7.1M/mo), pangu-py as pangu, and
strawberry-django as strawberry-graphql-django (1.5M/mo). httpx.URL is
left unmapped deliberately since it's a class within the httpx package,
not a package of its own.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 18:54:58 +08:00
Vinta ChenandClaude cc804b1de2 feat: add Downloads/Month column to website table
Sourced from website/data/pypi_downloads.tsv the same way
github_stars.json feeds the stars column. The new sortable column
sits between GitHub Stars and Last Commit on the homepage and
category pages, formatted with thousands separators like stars, with
an em dash when no PyPI data exists. Rows are matched by normalized
README display name; Built-in entries never show counts since
same-named PyPI packages are stdlib backports (e.g. the asyncio
package).

Below 960px the column hides and the count moves into the expand
row, mirroring the existing Last Commit treatment. main.js gains the
downloads sort branch and URL param.

The deploy workflow fetches the TSV via the new
make fetch_pypi_downloads target with a daily actions/cache
fallback, mirroring the stars fetch, but non-fatal: the column
degrades to dashes when the fetch fails, unlike stars which the
build requires.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 17:52:04 +08:00
Vinta ChenandClaude b440e64cd0 fix: resolve PyPI download counts through curated package overrides
A pypi.org identity sweep of all 438 cached rows (project_urls/home_page
vs entry GitHub URL) found download counts were looked up by README
display name, so entries whose name differs from the canonical package
silently measured squatters or dead predecessors: pytorch measured a
squatter (169,737/mo vs torch's 94M), jinja measured Jinja1 (3,168 vs
jinja2's 736M), django-rest-framework a dead alias package (real:
djangorestframework), django-rules an abandoned fork (real: rules),
strawberry an unrelated bookmarking service (real: strawberry-graphql),
devpi a deprecated metapackage (mapped to devpi-server).

New curated website/data/pypi_name_overrides.json maps normalized
README name to the real package, or null for projects not
pip-installable whose name is squatted or a relic (cpython, pyenv,
renpy, python-patterns, winpython); also maps mem0 to mem0ai, fasthtml
to python-fasthtml, and playwright-python to playwright.

All three fetch scripts resolve names through it; the clickpy TSV
cache gains a package column recording what each row actually
measured. .gitignore switches website/data/ to website/data/* with a
negation so the curated overrides file is tracked while caches stay
ignored.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 17:51:11 +08:00
Vinta ChenandClaude 2c2ce2fcac docs: document empirically verified README parsing behavior
Capture parser quirks worth knowing before editing README.md:
everything above  is ignored, new subcategories need no
parser change, a standalone all-bold paragraph becomes a Thematic
Group marker, prose after  leaks into llms.txt, and the
build's "Total entries" figure counts sub-items rather than just
entries.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:36 +08:00
Vinta ChenandClaude 2c50d9e63a docs: document gitignored output and orphan-key behavior
Note that data/github_stars.json is gitignored and fetched by CI at
deploy time, so local runs are preview-only and should never be
committed; entries removed from README.md just leave harmless orphan
keys behind.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:31 +08:00
Vinta ChenandClaude 39eeab7dff docs: warn about wrong-package PyPI download counts
The pypi downloads sweep looks up counts by README display name; when
the display name differs from the canonical package, the row silently
measures an unrelated squatter or a dead predecessor. Document the
failure mode in both the fetcher's docstring and the audit skill so
famous entries with off-looking counts get identity-verified before
being cited.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:15 +08:00
Vinta ChenandClaude a2da210345 feat: stamp pypi downloads cache with fetched_at and header row
Adds a header row (name, downloads, fetched_at) to data/pypi_downloads.tsv
and stamps every row with the sweep date, so audits can tell evidence age
and skip re-fetching when the cache is less than 7 days old. The sweep
itself costs ~1s, so freshness is checked by the reader (audit-the-list
skill) instead of skip logic in the fetch script. SKILL.md documents the
7-day freshness rule.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 02:58:20 +08:00
Vinta ChenandClaude 1abfb1af42 refactor: split download fetchers into writer and cross-check scripts
Authored by the parallel fetch-scripts session (its intended 0cffb5f
never landed; the content rode into an audit commit by accident and is
extracted here): fetch_pypi_downloads_via_clickpy.py becomes the
flagless full-README sweep and sole writer of data/pypi_downloads.tsv;
new fetch_pypi_downloads_via_bigquery.py is a print-only cross-check
taking explicit names behind a 400 GB billing cap; audit-the-list
SKILL.md step 2 updated to match.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:57:16 +08:00
Vinta ChenandClaude 5297061a4f refactor: rename download-fetch scripts to source-suffixed scheme
fetch_pypi_downloads.py becomes fetch_pypi_downloads_via_clickpy.py and
fetch_pepy_downloads.py becomes fetch_pypi_downloads_via_pepy.py, ahead
of splitting the BigQuery path into its own file. Updates the usage
strings, the cross-file import, and the audit-the-list skill's
references to match. Pure rename, no behavior change.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:35:24 +08:00
Vinta ChenandClaude 4c57da829d audit: sweep CMS to the shortlist cap
First Audit of the section; maintainer adjudicated 2026-08-16.
Reordered by downloads/month (wagtail, django-cms).

Removed (downloads are PyPI last-30-day via BigQuery, fetched 2026-08-16):

- indico (3.2K/month): event-management application (CERN), not a CMS
  library; readers under CMS want wagtail/django-cms - judgment call

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:34:07 +08:00
Vinta ChenandClaude fcac01955c feat: add pepy.tech spot-check helper for download audits
Maintainer registered a free pepy API key (PEPY_TECH_API_KEY in the gitignored repo-root .env). fetch_pepy_downloads.py sums the most recent 30 days from the v2 per-day data, throttled to the free tier's 5 requests/minute, and prints TSV without touching the single-source cache file.

The audit-the-list skill's downloads bullet is rewritten in the same commit because the ClickPy-default change made its old --dry-run-first instruction fail; it now documents all three sources (ClickPy, BigQuery, pepy/pypistats) and the never-mix-mirror-counting rule.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:19:47 +08:00
Vinta ChenandClaude fd79813dde feat: switch default PyPI download source to ClickPy
BigQuery required a personal Google Cloud account, so only the maintainer could run it and overruns cost real money. ClickPy is ClickHouse's free, keyless public mirror of the same PyPI download dataset, so it becomes the default source with a single batched query covering the full README in under a second.

BigQuery stays available behind a new --bigquery flag as the canonical-source cross-check; --dry-run now only applies to that path.

Live-tested: ClickPy full path (2/2 names), the --dry-run guard, and the BigQuery dry-run.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:17:59 +08:00
Vinta ChenandClaude a1ed550c2d fix: filter BigQuery PyPI download query on project, not file.project
Live bq show verified the pypi.file_downloads table clusters on the
top-level project column, not file.project as the docstring claimed.
Filtering on project (values verified identical to file.project across
408M rows, zero mismatches) gets cluster pruning and cuts the
full-README scan estimate from >1.2TB to ~275GB upper bound, with
actual billed bytes lower still (33.7GB measured for a single name) -
so full sweeps now fit the 1 TiB/month free tier.

Also adds --maximum_bytes_billed=400GB as a safety cap, enforced by
BigQuery pre-run against the dry-run upper-bound estimate.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:16:21 +08:00
Vinta ChenandClaude fc88ebb899 feat: add BigQuery-based PyPI downloads fetcher
Provides per-sitting download evidence for prune sweeps, per the
shortlist-reform tooling plan. Shells out to the bq CLI against
bigquery-public-data.pypi.file_downloads, parses entry names from
README.md via readme_parser, and supports --dry-run and --names-file.
Merges results into the gitignored cache at
website/data/pypi_downloads.tsv.

The table is clustered on file.project, so scanned bytes grow with the
IN-list size: a dry run against the full README (~530 names) scanned
1.21 TB, past the 1 TB/month free tier. Per-sitting --names-file
fetches are used instead of one big query.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-15 15:41:55 +08:00
Vinta Chen c0b80fd75f fix: surface sponsorship link in header
Deploy Website / deploy (push) Has been cancelled
CI / test (push) Has been cancelled
2026-06-07 03:22:28 +08:00
Vinta Chen 32ae78fd96 fix: improve website SEO metadata 2026-06-07 03:08:43 +08:00
Vinta Chen ee08cd7d86 update sponsorship webpage 2026-06-07 02:50:19 +08:00
Vinta Chen 5f725c25d7 use sponsorship@awesome-python.com as contact
Deploy Website / deploy (push) Has been cancelled
CI / test (push) Has been cancelled
2026-05-07 21:14:26 +08:00
Vinta Chen 10c06fb26d add category links in llms.txt 2026-05-07 20:00:56 +08:00
Vinta Chen 6c18b6447e feat: use explicit Projects section in README
Deploy Website / deploy (push) Has been cancelled
CI / test (push) Has been cancelled
2026-05-04 21:24:57 +08:00
Vinta Chen 921d47b455 remove index.md 2026-05-04 17:11:34 +08:00
Vinta Chen 3510db9df9 update llms.txt 2026-05-04 17:05:05 +08:00
Vinta Chen 509ebaff7a use file modification time as lastmod in sitemap 2026-05-04 16:24:52 +08:00
Vinta ChenandClaude 9379e0a42c style(sitemap): pretty-print generated sitemap.xml with 2-space indent
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 20:16:13 +08:00
Vinta ChenandClaude 28b61a9212 style(seo): switch category page title separator from pipe to hyphen
Google truncates pipe separators and treats hyphens as cleaner word
boundaries in SERP titles.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 20:03:29 +08:00
Vinta ChenandClaude c886e470b6 feat(website): lead category meta description with real description when present, count first as fallback
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 19:57:31 +08:00
Vinta ChenandClaude 2f398acefb fix(seo): align JSON-LD with Yoast/RankMath conventions
- Wrap category pages in a self-contained @graph (WebSite + CollectionPage)
- Set canonical @id on CollectionPage to its URL (no hash fragment)
- Expand isPartOf to typed object {"@type": "WebSite", "@id": ...}
- Extract _website_node() and ISPARTOF_WEBSITE constants to avoid repetition
- Update tests to assert @graph structure on category pages

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 19:31:14 +08:00
Vinta ChenandClaude 86d2aa7e01 feat(website): add CollectionPage JSON-LD to category, group, and subcategory pages
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 19:23:14 +08:00
Vinta ChenandClaude b2910d59c8 feat(website): add homepage JSON-LD with WebSite, CollectionPage, ItemList for SEO/AEO
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 19:18:15 +08:00
Vinta ChenandClaude 138059feeb feat(website): add Awesome Python and Sponsorship links to footer
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 19:05:40 +08:00
Vinta ChenandClaude f57fc44295 style: bump tag font size to var(--text-xs) and codify 12px minimum font-size rule
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 18:54:23 +08:00
Vinta ChenandClaude f3f92c691a feat(website): render subcategory, group, and source tags as anchor elements
Convert <button> tags for subcategory, group, and source filters to <a>
elements with href attributes so browsers surface URL preview on hover,
support open-in-new-tab, and allow middle-click navigation.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 18:49:34 +08:00
Vinta ChenandClaude 9de86ea785 feat(website): append #library-index to tag links on non-index pages
Tag clicks on category/other pages now land with the results section
scrolled into view.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 18:45:09 +08:00
Vinta ChenandClaude fc8d1ba35e feat(website): show desc-row on index page when a filter is active
On category pages desc-rows are always visible. On the index page they
were always hidden. Now they become visible whenever a tag/category
filter is applied, giving filtered results the same richness as category
pages.

Also tightens two related CSS rules: border-bottom suppression only
fires when the adjacent desc-row is actually visible, and the expand-row
description is hidden while the desc-row is already showing to avoid
duplicate text.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 13:08:27 +08:00
Vinta ChenandClaude 1468ae78ff feat(website): show project description as always-visible desc-row on category pages
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 12:47:17 +08:00
Vinta ChenandClaude 3d99f7336d style(website): apply ruff format
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 12:23:55 +08:00
Vinta ChenandClaude d3f35a9d21 test(website): remove redundant and brittle tests
Drops tests that either duplicate coverage already provided by adjacent
cases (single-word slugify, trailing-slash checks) or hard-code first-
category names and specific description strings that break whenever the
README content shifts.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 12:19:32 +08:00
Vinta Chen a068219684 fix(website): type build template entries 2026-05-03 12:08:41 +08:00
Vinta ChenandClaude 38b54caabb fix(website): trim sponsorship page nav and hero stats
Remove the 'All projects' nav link and total_entries hero stat from the
sponsorship page. Rename 'View the repository' CTA to 'View on GitHub'.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 11:45:07 +08:00
Vinta ChenandClaude ee01a0bade refactor(website): extract render_category, replace slugify filter with filter_urls map
- Extract render_category() helper to deduplicate the three category/group/builtin
  rendering blocks in build.py
- Replace synthetic dict literals with synthetic_category() helper
- Rewrite subcategory rendering to avoid O(n²) loop using precomputed dicts
- Pass filter_urls (not just JSON) to templates so Jinja can look up group URLs
  directly instead of applying the slugify filter at render time
- Remove slugify from env.filters (no longer used in templates)
- Replace isIndexPage() wrapper with isIndexDocument constant in main.js
- Fix: call applyFilters() on page load when activeFilter is set
- Remove dead else branch in tag click handler (category pages with no URL)
- Switch .hero-category-links from CSS columns to CSS grid for more even layout
- Remove max-width cap on .category-subtitle

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 11:38:22 +08:00
Vinta ChenandClaude 8a32d27ef5 fix(website): tighten sponsorship page copy for clarity
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 09:46:50 +08:00
Vinta Chen 40913c3df3 fix double quotes 2026-05-03 09:38:28 +08:00
Vinta ChenandClaude c68b985d7c feat(website): add /sponsorship/ landing page
Adds a dedicated sponsorship page at /sponsorship/ built from the Jinja2
template, with hero stats, tier cards, and CSS. Updates the index.html
sponsor sidebar link to point to /sponsorship/ instead of the GitHub
SPONSORSHIP.md. Adds the URL to the sitemap and test fixtures.

Also renames .impeccable.md to DESIGN.md.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 09:35:39 +08:00
Vinta ChenandClaude 64781112d8 feat(website): add Browse by category nav to group page hero
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 09:08:37 +08:00
Vinta ChenandClaude b82a254a09 fix(website): clear filter lands at /#library-index on category pages
Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 09:03:25 +08:00
Vinta ChenandClaude 70a8255289 feat(website): add /categories/built-in/ page for Built-in tag filter
Register Built-in as a navigable filter path alongside regular category
and group slugs, emit the page during build, add it to the sitemap, and
wire the Built-in tag buttons in index.html and category.html to navigate
there via data-url.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-05-03 08:35:55 +08:00