Files
awesome-python/website
Vinta ChenandClaude a1ed550c2d fix: filter BigQuery PyPI download query on project, not file.project
Live bq show verified the pypi.file_downloads table clusters on the
top-level project column, not file.project as the docstring claimed.
Filtering on project (values verified identical to file.project across
408M rows, zero mismatches) gets cluster pruning and cuts the
full-README scan estimate from >1.2TB to ~275GB upper bound, with
actual billed bytes lower still (33.7GB measured for a single name) -
so full sweeps now fit the 1 TiB/month free tier.

Also adds --maximum_bytes_billed=400GB as a safety cap, enforced by
BigQuery pre-run against the dry-run upper-bound estimate.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 01:16:21 +08:00
..
2026-06-07 03:08:43 +08:00