mirror of
https://github.com/vinta/awesome-python.git
synced 2026-10-02 08:23:10 +08:00
Live bq show verified the pypi.file_downloads table clusters on the top-level project column, not file.project as the docstring claimed. Filtering on project (values verified identical to file.project across 408M rows, zero mismatches) gets cluster pruning and cuts the full-README scan estimate from >1.2TB to ~275GB upper bound, with actual billed bytes lower still (33.7GB measured for a single name) - so full sweeps now fit the 1 TiB/month free tier. Also adds --maximum_bytes_billed=400GB as a safety cap, enforced by BigQuery pre-run against the dry-run upper-bound estimate. Co-Authored-By: Claude <noreply@anthropic.com>