Adds a header row (name, downloads, fetched_at) to data/pypi_downloads.tsv
and stamps every row with the sweep date, so audits can tell evidence age
and skip re-fetching when the cache is less than 7 days old. The sweep
itself costs ~1s, so freshness is checked by the reader (audit-the-list
skill) instead of skip logic in the fetch script. SKILL.md documents the
7-day freshness rule.
Co-Authored-By: Claude <noreply@anthropic.com>
Authored by the parallel fetch-scripts session (its intended 0cffb5f
never landed; the content rode into an audit commit by accident and is
extracted here): fetch_pypi_downloads_via_clickpy.py becomes the
flagless full-README sweep and sole writer of data/pypi_downloads.tsv;
new fetch_pypi_downloads_via_bigquery.py is a print-only cross-check
taking explicit names behind a 400 GB billing cap; audit-the-list
SKILL.md step 2 updated to match.
Co-Authored-By: Claude <noreply@anthropic.com>
fetch_pypi_downloads.py becomes fetch_pypi_downloads_via_clickpy.py and
fetch_pepy_downloads.py becomes fetch_pypi_downloads_via_pepy.py, ahead
of splitting the BigQuery path into its own file. Updates the usage
strings, the cross-file import, and the audit-the-list skill's
references to match. Pure rename, no behavior change.
Co-Authored-By: Claude <noreply@anthropic.com>
First Audit of the section; maintainer adjudicated 2026-08-16.
Reordered by downloads/month (wagtail, django-cms).
Removed (downloads are PyPI last-30-day via BigQuery, fetched 2026-08-16):
- indico (3.2K/month): event-management application (CERN), not a CMS
library; readers under CMS want wagtail/django-cms - judgment call
Co-Authored-By: Claude <noreply@anthropic.com>
Maintainer registered a free pepy API key (PEPY_TECH_API_KEY in the gitignored repo-root .env). fetch_pepy_downloads.py sums the most recent 30 days from the v2 per-day data, throttled to the free tier's 5 requests/minute, and prints TSV without touching the single-source cache file.
The audit-the-list skill's downloads bullet is rewritten in the same commit because the ClickPy-default change made its old --dry-run-first instruction fail; it now documents all three sources (ClickPy, BigQuery, pepy/pypistats) and the never-mix-mirror-counting rule.
Co-Authored-By: Claude <noreply@anthropic.com>
BigQuery required a personal Google Cloud account, so only the maintainer could run it and overruns cost real money. ClickPy is ClickHouse's free, keyless public mirror of the same PyPI download dataset, so it becomes the default source with a single batched query covering the full README in under a second.
BigQuery stays available behind a new --bigquery flag as the canonical-source cross-check; --dry-run now only applies to that path.
Live-tested: ClickPy full path (2/2 names), the --dry-run guard, and the BigQuery dry-run.
Co-Authored-By: Claude <noreply@anthropic.com>
Live bq show verified the pypi.file_downloads table clusters on the
top-level project column, not file.project as the docstring claimed.
Filtering on project (values verified identical to file.project across
408M rows, zero mismatches) gets cluster pruning and cuts the
full-README scan estimate from >1.2TB to ~275GB upper bound, with
actual billed bytes lower still (33.7GB measured for a single name) -
so full sweeps now fit the 1 TiB/month free tier.
Also adds --maximum_bytes_billed=400GB as a safety cap, enforced by
BigQuery pre-run against the dry-run upper-bound estimate.
Co-Authored-By: Claude <noreply@anthropic.com>
Provides per-sitting download evidence for prune sweeps, per the
shortlist-reform tooling plan. Shells out to the bq CLI against
bigquery-public-data.pypi.file_downloads, parses entry names from
README.md via readme_parser, and supports --dry-run and --names-file.
Merges results into the gitignored cache at
website/data/pypi_downloads.tsv.
The table is clustered on file.project, so scanned bytes grow with the
IN-list size: a dry run against the full README (~530 names) scanned
1.21 TB, past the 1 TB/month free tier. Per-sitting --names-file
fetches are used instead of one big query.
Co-Authored-By: Claude <noreply@anthropic.com>
- Wrap category pages in a self-contained @graph (WebSite + CollectionPage)
- Set canonical @id on CollectionPage to its URL (no hash fragment)
- Expand isPartOf to typed object {"@type": "WebSite", "@id": ...}
- Extract _website_node() and ISPARTOF_WEBSITE constants to avoid repetition
- Update tests to assert @graph structure on category pages
Co-Authored-By: Claude <noreply@anthropic.com>
Convert <button> tags for subcategory, group, and source filters to <a>
elements with href attributes so browsers surface URL preview on hover,
support open-in-new-tab, and allow middle-click navigation.
Co-Authored-By: Claude <noreply@anthropic.com>
On category pages desc-rows are always visible. On the index page they
were always hidden. Now they become visible whenever a tag/category
filter is applied, giving filtered results the same richness as category
pages.
Also tightens two related CSS rules: border-bottom suppression only
fires when the adjacent desc-row is actually visible, and the expand-row
description is hidden while the desc-row is already showing to avoid
duplicate text.
Co-Authored-By: Claude <noreply@anthropic.com>
Drops tests that either duplicate coverage already provided by adjacent
cases (single-word slugify, trailing-slash checks) or hard-code first-
category names and specific description strings that break whenever the
README content shifts.
Co-Authored-By: Claude <noreply@anthropic.com>
Remove the 'All projects' nav link and total_entries hero stat from the
sponsorship page. Rename 'View the repository' CTA to 'View on GitHub'.
Co-Authored-By: Claude <noreply@anthropic.com>
- Extract render_category() helper to deduplicate the three category/group/builtin
rendering blocks in build.py
- Replace synthetic dict literals with synthetic_category() helper
- Rewrite subcategory rendering to avoid O(n²) loop using precomputed dicts
- Pass filter_urls (not just JSON) to templates so Jinja can look up group URLs
directly instead of applying the slugify filter at render time
- Remove slugify from env.filters (no longer used in templates)
- Replace isIndexPage() wrapper with isIndexDocument constant in main.js
- Fix: call applyFilters() on page load when activeFilter is set
- Remove dead else branch in tag click handler (category pages with no URL)
- Switch .hero-category-links from CSS columns to CSS grid for more even layout
- Remove max-width cap on .category-subtitle
Co-Authored-By: Claude <noreply@anthropic.com>
Adds a dedicated sponsorship page at /sponsorship/ built from the Jinja2
template, with hero stats, tier cards, and CSS. Updates the index.html
sponsor sidebar link to point to /sponsorship/ instead of the GitHub
SPONSORSHIP.md. Adds the URL to the sitemap and test fixtures.
Also renames .impeccable.md to DESIGN.md.
Co-Authored-By: Claude <noreply@anthropic.com>
Register Built-in as a navigable filter path alongside regular category
and group slugs, emit the page during build, add it to the sitemap, and
wire the Built-in tag buttons in index.html and category.html to navigate
there via data-url.
Co-Authored-By: Claude <noreply@anthropic.com>
Add search input, filter chips, no-results block, and back-to-top
button to category/group/subcategory pages. Pass filter_urls_json to
all page types so tag-chip navigation works site-wide. Fix JS so
filter-clear and no-results-clear redirect to / on non-index pages
instead of trying to filter a non-existent local table. Remove the
now-redundant .category-results CSS overrides.
Co-Authored-By: Claude <noreply@anthropic.com>
Removes inline .category-row-desc from the name cell and renders
entry.description inside .expand-content instead, matching the
index page pattern. Drops the now-unused CSS rules for
.category-row-desc and the overridden .category-table .expand-content
padding.
Co-Authored-By: Claude <noreply@anthropic.com>
The results-intro grid (1fr + 28rem note column) squeezed the heading on
category pages with long names, e.g. "Python Projects in Environment
Management" wrapped onto two lines.
Scope a single-column override to .category-results so the heading takes
the full row and the note drops below right-aligned. Index page layout
is untouched since its heading is short.
The "All projects" link in the category-page topbar pointed to
/#library-index so the browser would scroll to the library section on
arrival. The hash stayed in the URL, which looked like an internal anchor
state rather than a clean homepage URL.
On homepage load, if the hash is #library-index, scroll to the section
explicitly and use history.replaceState to drop the hash from the URL.
The scrollIntoView call covers the case where the script runs before the
browser's native anchor scroll, since replaceState removes the hash the
browser would have used.
The category template rendered a tag for `category.name` plus a tag for
`entry.groups[0]`, which duplicated the group name on group pages where
those values are identical (e.g. /categories/python-language/ showing
"Python Language" twice). It also never rendered `entry.categories`, so
group pages omitted each project's actual category.
Mirror the index template's tag rendering on category, group, and
subcategory pages, and mark whichever tag matches the current page URL
as active. Pass `category_urls` and `current_path` to each render call
so the template can match by URL.
Tag clicks on / pushState a category/group/subcategory path; on static
pages they fully navigate. Search and sort stay in querystring. Built-in
source tag has no data-url and stays as an in-page filter. The
isIndexDocument flag is captured at load time so toggling on the index
keeps working after pushState changes location.pathname.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
`| safe` bypasses Jinja autoescape. If a category name ever contained
"</script>", the literal substring would close the script block early,
leaking JSON content into the DOM and creating an XSS vector. Replace
"</" with "<\\/" (still valid JSON) and pass ensure_ascii=False so
non-ASCII names render readably. Also add a group_path() helper to
parallel category_path()/subcategory_path() and reuse category_urls
when seeding filter_urls.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds filter_urls dict (categories, groups, subcategories) in build.py,
passes filter_urls_json to the template, and injects a JSON script block
before the results section in index.html. Covered by a new test that
verifies all three URL types are present and correctly resolved.
Co-Authored-By: Claude <noreply@anthropic.com>