Commit Graph
2685 Commits
Author SHA1 Message Date
Vinta ChenandClaude 4fc35351b2 docs: update og-image wordings to match current site copy
Kicker now mirrors the hero kicker ("The definitive list that
answers..."), replacing the old "field guide" line changed in
8b14b4f. Subtitle now matches the current tagline ("An opinionated
guide to the best Python frameworks, libraries, and tools.") on one
line. PNG regenerated from the SVG.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 20:31:39 +08:00
Vinta Chen 8b14b4fd1e update styles 2026-08-16 20:25:19 +08:00
Vinta Chen e63de26416 update readme 2026-08-16 20:19:01 +08:00
Vinta Chen ee8089d0d5 add build date to llms.txt 2026-08-16 20:16:05 +08:00
Vinta ChenandClaude 30c9f18bc2 feat: annotate llms.txt entries with PyPI download counts
Entries now end with "(PyPI downloads/month: N, GitHub stars: M)" where
known, replacing the stars-only note, since download counts are the
list's stated primary evidence signal. annotate_entries_with_stars is
renamed to annotate_entries_with_stats and looks downloads up by the
first link's display name, skipping category-index bullets (which link
into the site itself) and Built-in entries (which would otherwise hit
same-named PyPI backports like logging or asyncio).

The intro now mirrors the README subtitle verbatim with the
project/category totals on their own line below, and the "opinionated
catalog" wording is gone since the shortlist ADR is literally titled
"shortlist, not a catalog".

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:43:52 +08:00
Vinta Chen b7e12cfcad Merge pull request #3288 from vinta/refactor/shortlist-reform
Reposition awesome-python as a shortlist, not a catalog
2026-08-16 19:31:51 +08:00
Vinta ChenandClaude bfdd087282 feat: show "Not on PyPI" badge for missing download counts
Replaces the em dash in the PyPI Downloads column with a source-badge
pill labeled "Not on PyPI", reusing the existing badge style used by
the stars column for visual consistency. Sorting is unaffected since
non-numeric cells already parse as missing.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:20:55 +08:00
Vinta ChenandClaude f5f7a1c50c refactor: swap PyPI Downloads and GitHub Stars column order
Downloads is now the default sort, so it sits directly after the
project name in both the index and category table templates. The
source-type badge stays in the stars cell.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:19:42 +08:00
Vinta ChenandClaude da491ff8cc docs: rename Downloads/Month column header to PyPI Downloads
The header no longer carries the per-month unit; expand-row text keeps
its 'downloads/month' wording.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:16:03 +08:00
Vinta ChenandClaude d3fec1f7fa feat: default-sort the website table by downloads per month
Entries with a download count now sort first (descending), with
stars, then Built-in, then name as fallback tiers for entries that
lack a count. main.js mirrors this in its default activeSort, clean
URL check, and third-click reset target. Sorting by stars remains one
header click away.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:15:21 +08:00
Vinta ChenandClaude b841db68ba refactor: make pypi_name_overrides.json entries self-documenting
Every entry is now {"package": str|null, "reason": str|null} instead of
a bare string/null. Reasons are required for null packages, explaining
why the name must never be queried (squatted name, stdlib module,
monorepo umbrella, GitHub-only project, and so on). Reasons are
optional for remaps and kept only on the six non-obvious ones: pytorch
(squatter), jinja (jinja is Jinja1), strawberry (unrelated bookmarking
service), django-rules (abandoned fork), django-rest-framework (dead
alias), and devpi (deprecated metapackage); plain publishes-as-X
remaps get a null reason.

load_overrides() in the clickpy fetcher now extracts the package field
from each entry; resolve() and the pepy/bigquery cross-check scripts
are unchanged since they consume load_overrides()'s output.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 19:08:55 +08:00
Vinta ChenandClaude fcfc65df00 data: record all remaining NOT_FOUND names as explicit null overrides
Every queried name now resolves 447/447. Adds 23 explicit null
overrides so squatters can never silently attach a PyPI number to
these names later: stdlib-named entries (concurrent-futures, difflib,
mimetypes, sqlite3, tkinter, tomllib, zoneinfo), interpreters
(micropython, pypy), monorepo umbrellas (azure-sdk-for-python,
google-cloud-python), self-hosted or distro-installed projects (odoo,
cloud-init, warehouse), GitHub-only projects (thealgorithms,
geodjango, django-db-models, django-ai-plugins, graphify,
sentry-skills, social-engineer-toolkit, trailofbits-skills), and
httpx-url (a class within httpx, not a package).

Caveat: graphify and django-ai-plugins are young projects that may
legitimately publish to PyPI later — flip their null to a remap
during a future audit if they do.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 18:57:13 +08:00
Vinta ChenandClaude f3ff733fd7 fix: add three more PyPI name remaps for NOT_FOUND triage
autobahn-python publishes as autobahn (7.1M/mo), pangu-py as pangu, and
strawberry-django as strawberry-graphql-django (1.5M/mo). httpx.URL is
left unmapped deliberately since it's a class within the httpx package,
not a package of its own.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 18:54:58 +08:00
Vinta ChenandClaude cc804b1de2 feat: add Downloads/Month column to website table
Sourced from website/data/pypi_downloads.tsv the same way
github_stars.json feeds the stars column. The new sortable column
sits between GitHub Stars and Last Commit on the homepage and
category pages, formatted with thousands separators like stars, with
an em dash when no PyPI data exists. Rows are matched by normalized
README display name; Built-in entries never show counts since
same-named PyPI packages are stdlib backports (e.g. the asyncio
package).

Below 960px the column hides and the count moves into the expand
row, mirroring the existing Last Commit treatment. main.js gains the
downloads sort branch and URL param.

The deploy workflow fetches the TSV via the new
make fetch_pypi_downloads target with a daily actions/cache
fallback, mirroring the stars fetch, but non-fatal: the column
degrades to dashes when the fetch fails, unlike stars which the
build requires.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 17:52:04 +08:00
Vinta ChenandClaude b440e64cd0 fix: resolve PyPI download counts through curated package overrides
A pypi.org identity sweep of all 438 cached rows (project_urls/home_page
vs entry GitHub URL) found download counts were looked up by README
display name, so entries whose name differs from the canonical package
silently measured squatters or dead predecessors: pytorch measured a
squatter (169,737/mo vs torch's 94M), jinja measured Jinja1 (3,168 vs
jinja2's 736M), django-rest-framework a dead alias package (real:
djangorestframework), django-rules an abandoned fork (real: rules),
strawberry an unrelated bookmarking service (real: strawberry-graphql),
devpi a deprecated metapackage (mapped to devpi-server).

New curated website/data/pypi_name_overrides.json maps normalized
README name to the real package, or null for projects not
pip-installable whose name is squatted or a relic (cpython, pyenv,
renpy, python-patterns, winpython); also maps mem0 to mem0ai, fasthtml
to python-fasthtml, and playwright-python to playwright.

All three fetch scripts resolve names through it; the clickpy TSV
cache gains a package column recording what each row actually
measured. .gitignore switches website/data/ to website/data/* with a
negation so the curated overrides file is tracked while caches stay
ignored.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 17:51:11 +08:00
Vinta ChenandClaude e9329d4d1e fix: correct hydra-core repo URL to facebookresearch/hydra
The entry linked hydra-ecosystem/hydra, an unrelated W3C Hydra API
toolkit, while the entry name and description describe
facebookresearch's Hydra configuration framework, mixing the wrong
repo's stars with the right package's identity. Found during the
downloads-column identity sweep.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 17:50:33 +08:00
Vinta ChenandClaude 0d13cd27f1 build: use >= floors instead of == pins for dependency groups
Exact reproducibility already lives in uv.lock via 'uv sync --locked',
so == in pyproject.toml only duplicates the lockfile and blocks
'uv lock --upgrade'. Locked versions are unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:59:09 +08:00
Vinta ChenandClaude 24d6c2896b build: upgrade project to Python 3.14
watchdog 6.0.0 (last release 2024-11-01) ships no cp314 macOS wheel,
and uv has no per-package build allowlist under no-build = true, so
the preview file watcher moves to watchfiles, which ships cp314
wheels. watchfiles now lives in its own preview dependency group.
UV_PYTHON=3.13 is no longer needed on machines that only have 3.14.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:58:34 +08:00
Vinta ChenandClaude c39e4d6bf5 docs: restrict sub-items to awesome-* also-see links
Removed the companion-project clause from the Sub-item definition in
CONTEXT.md's vocabulary. Its examples (aws-sdk-pandas under pandas,
flower under celery) went stale this sitting: those companions were
promoted, re-homed, or deleted. Per the maintainer's 2026-08-16 policy
decision, sub-items are now reserved for awesome-* also-see links only
- a companion project must earn a full Entry in its proper Use Case or
not be listed.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:47:34 +08:00
Vinta ChenandClaude ecd9dea9e4 refactor: move pyenv-win from pyenv sub-item to full entry
Re-homed pyenv-win from a pyenv sub-item (Environment Management) to a full entry in Microsoft Windows, placed before winpython by downloads (25.8k/mo vs 172). Actively maintained, pushed 2026-08-14, 7,360 stars. Maintainer preference is to move sub-items to a fitting category rather than delete.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:47:02 +08:00
Vinta ChenandClaude 49b229bab4 refactor: promote flower from celery sub-item to full DevOps Monitoring entry
Flower isn't a task queue, so nesting it under celery misclassified it; Task Queues is also at its entry cap. Monitoring and Processes is its honest home, ranking fourth by downloads (12.35M/mo ClickPy, between supervisor 17.0M and sh 11.8M), and Celery's own docs name it the recommended monitor. Repo pushed 2026-08-16 with 7,232 stars. This fills Monitoring and Processes to its 5-entry cap.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:46:31 +08:00
Vinta ChenandClaude b0434abae6 refactor: promote mkdocs-material to full Documentation entry
Was a sub-item under mkdocs. By downloads it ranks second in the
section at 17.6M/mo (ClickPy), above mkdocs' 17.4M, and it powers
FastAPI, Pydantic, and Ruff/Polars docs (27,269 stars, pushed
2026-08-09). Documentation now sits at its 5-entry cap.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:45:57 +08:00
Vinta ChenandClaude 63a1eadd30 docs: remove typeshed sub-item from under mypy
Not a tool readers install: type checkers bundle it automatically as a
stub collection, it has no PyPI package, and no standalone use case.
The Type Checkers subcategory label already links to
awesome-python-typing for ecosystem depth.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:45:26 +08:00
Vinta ChenandClaude f13075b2af refactor: re-home aws-sdk-pandas from pandas sub-item to ETL General
Sub-item policy reserves sub-items for awesome-* links. aws-sdk-pandas
promoted out as awswrangler in Data Ingestion / ETL > General
(85.3M downloads/mo, 10x dlt, active).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:44:38 +08:00
Vinta Chen 0e9ae0cb04 clean up 2026-08-16 16:25:10 +08:00
Vinta Chen 08fdfa6a88 update readme 2026-08-16 16:25:00 +08:00
Vinta ChenandClaude 68b15a4644 docs: rephrase obvious-choice criterion in PR template
Reuses CONTRIBUTING.md's plainer 'would name when asked' phrasing instead of 'unprompted', per maintainer feedback that 'unprompted' didn't sound right.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:13:43 +08:00
Vinta ChenandClaude cccfa183a9 docs: rewrite PR template to match shortlist rules
The old template used a stars-based tier system (Industry Standard /
Rising Star / Hidden Gem) that contradicted the current CONTRIBUTING.md,
which judges entries by obvious-choice/challenger tiers, favors PyPI
downloads over stars, and requires Displacement when a use case is at
its cap. The new template reflects those rules and adds a checklist
item pointing contributors to CONTRIBUTING.md.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:11:00 +08:00
Vinta ChenandClaude cfb60a633f docs: codify dual-listing policy in CONTRIBUTING.md
The maintainer decided how duplicate entries across categories should
be handled (e.g. uv listed in both Environment Management and Package
Management): each slot must earn its place independently, entries are
listed in full with identical lines rather than a cross-reference,
description edits update every copy in the same commit, and each slot
is audited on its own.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 16:06:34 +08:00
Vinta ChenandClaude 2c2ce2fcac docs: document empirically verified README parsing behavior
Capture parser quirks worth knowing before editing README.md:
everything above  is ignored, new subcategories need no
parser change, a standalone all-bold paragraph becomes a Thematic
Group marker, prose after  leaks into llms.txt, and the
build's "Total entries" figure counts sub-items rather than just
entries.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:36 +08:00
Vinta ChenandClaude 2c50d9e63a docs: document gitignored output and orphan-key behavior
Note that data/github_stars.json is gitignored and fetched by CI at
deploy time, so local runs are preview-only and should never be
committed; entries removed from README.md just leave harmless orphan
keys behind.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:31 +08:00
Vinta ChenandClaude deae0e0b72 docs: broaden Obvious Choice signal and note Cap has no floor
Extend the known failure-mode list for PyPI download counts beyond
model weights to any project consumed outside pip (SDK downloads like
renpy, deployed services like thumbor). Also clarify that the per-Use-
Case Cap is a ceiling, not a floor: a freshly minted Use Case may hold
a single entry.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:29 +08:00
Vinta ChenandClaude 3e43fa13ba docs: clarify cross-section re-homes and resources scope in audit rules
Cross-section re-homes now ride the originating audit's commit instead
of needing a separate one, since both sides of the move land in one
diff. Also note that Resources sections are out of audit scope and
never parsed by the website, so they're not project entries subject
to the one-entry-per-commit rule.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:25 +08:00
Vinta ChenandClaude c034c88e91 docs: add audit log registry for maintainer overrides
Git history already archives every removal's reason via commit body,
but it can't be scanned at a glance. docs/audit-logs.md is the
at-a-glance register of overrides (naming exceptions, mature-stable
keeps) allowed by CONTRIBUTING.md. Drop the docs/* gitignore exclusion
(and stale .superpowers/ and skills-lock.json entries) so the file and
future doc additions outside docs/adr/ can be tracked.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:21 +08:00
Vinta ChenandClaude 39eeab7dff docs: warn about wrong-package PyPI download counts
The pypi downloads sweep looks up counts by README display name; when
the display name differs from the canonical package, the row silently
measures an unrelated squatter or a dead predecessor. Document the
failure mode in both the fetcher's docstring and the audit skill so
famous entries with off-looking counts get identity-verified before
being cited.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:57:15 +08:00
Vinta ChenandClaude 194b386d06 docs: codify the Override rule in CONTRIBUTING and CONTEXT
Resolves decision 10's reservation on maintainer word: the 3+2/5 cap
numbers stay as written — across the full prune they held everywhere
except a handful of explicit overrides — and the override practice
itself becomes a written rule: the maintainer may exceed any limit
for a specific entry or use case by explicit decision, case-by-case,
carrying no weight for submissions. CONTEXT.md gains the matching
Override vocabulary entry.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:25:15 +08:00
Vinta ChenandClaude f60d5b4a08 style: move fasthtml back to Web Frameworks > Synchronous
Maintainer reversal of c0a31ce, restoring the entry and its
awesome-fasthtml sub-item to their prior position. The
challenger-limit override that rode the move is withdrawn with it —
Asynchronous returns to 4 entries within the standard cap shape.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:25:15 +08:00
Vinta ChenandClaude 441f206d71 feat: add pathway to Data Ingestion / ETL
Re-admission on maintainer word, reversing the Data Analysis sweep's
drop (f3c920d — the xlsxwriter reversal precedent): the drop was
partly a mis-homing casualty, since its honest home, an ETL use case,
did not exist then. The repo self-describes as a Python ETL framework
for stream processing and LLM/RAG pipelines: 62.5K stars, pushed
daily; 16.5K downloads/month is weak for the star count and noted.
Enters as challenger behind dlt (7.8M/mo).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:24:38 +08:00
Vinta ChenandClaude e5f7b5bc2e style: update graphify link to the moved Graphify-Labs repo
The safishamsi/graphify URL is a stale redirect — the repo moved to
Graphify-Labs/graphify (verified via the GitHub API). Maintainer
declined the Agent Skills re-home; the entry stays in Data
Visualization > Specialized with its link fixed.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:21:21 +08:00
Vinta ChenandClaude 5249b756ce refactor: re-home fasthtml to Web Frameworks > Asynchronous
Maintainer-adjudicated close of the standing flag: fasthtml runs on
Starlette and Uvicorn (ASGI), so Synchronous was the wrong shelf. It
lands as a third challenger behind starlette and tornado's obvious
choices — the use case holds 5 with 3 challengers by explicit
maintainer override (Async I/O precedent; the awesome-fasthtml
sub-item rides along). 1.19M downloads/month as python-fasthtml.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:21:08 +08:00
Vinta ChenandClaude 15cd51c04c style: re-tier File Manipulation
No removals. mimetypes and pathlib (standard library) lead
alphabetically under the stdlib-first rule; watchfiles (389.3M/mo,
partly uvicorn-transitive — the riser) completes the obvious choices;
watchdog (113.3M/mo, the demoted incumbent, Second Tier) and
python-magic (32.3M/mo) challengers.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:10:10 +08:00
Vinta ChenandClaude ad698c2876 audit: sweep HTML Manipulation, drop html-to-markdown, pyquery, tinycss2
Tiers: beautifulsoup4 (renamed from beautifulsoup — the bare PyPI
name is the abandoned bs3 shim; 451.4M/mo, docs link per the PyQt
precedent), lxml (401.3M/mo), xmltodict (124.5M/mo) obvious choices;
markupsafe (820.5M/mo — the section's biggest raw count, but
jinja-transitive infrastructure, so challenger on judgment; watch:
quiet since 2025-09) and justhtml (67.8K/mo, 1.1K stars in two
years — trajectory judgment on a young pure-Python HTML5 parser)
challengers.

Removed:
- html-to-markdown — coordinated multi-entry self-promotion
  (automatic-rejection rule): PyPI provenance verified to xberg-io,
  the org's fourth planted entry overall. 1.5M downloads/month is
  real but the rule stands.
- pyquery — 2.2M downloads/month and an active repo (pushed
  2026-07); editorial drop at cap: the jQuery-style API is the
  least-reached-for of the keeps. Judgment call.
- tinycss2 — 110.5M downloads/month is transitive (weasyprint
  declares it a hard dependency, verified in PyPI metadata) against
  190 stars; a CSS parser mis-homed in an HTML/XML section with no
  better home. Judgment call.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:09:52 +08:00
Vinta ChenandClaude db9c262342 feat: add rapidfuzz, minting Text Processing > Fuzzy Matching
The industry's fuzzy string matching answer (web-verified: the
production recommendation over thefuzz — same API, MIT license, C++
speed — and preferred over textdistance for string metrics). 181.7M
downloads/month (pepy), 4.1K stars, pushed 2026-08. Sole obvious
choice; the subcategory label rides this commit so it is never empty,
completing the displacement of textdistance.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:09:32 +08:00
Vinta ChenandClaude 043cb6d377 feat: add charset-normalizer to Text Processing > Encoding and Unicode
The ecosystem's default encoding detector — requests switched to it
in 2021 and 2026 guidance names it the choice for new projects
(web-verified). 1.73B downloads/month (pepy; heavily
requests-transitive, but default-status is the point), pushed
2026-08. Co-obvious with chardet, which retains a verified accuracy
claim — the PyQt/PySide pair shape.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:09:18 +08:00
Vinta ChenandClaude da4383a06f audit: sweep Text Processing, dissolve General, drop textdistance, nameparser, user-agents, tree-sitter-language-pack
Restructure: the 10-entry General grab-bag dissolves — Encoding and
Unicode (chardet 224.1M/mo obvious choice, joined by
charset-normalizer next commit; ftfy 14.4M/mo kept as the fifth
mature-stable past-line keep, repo and release both 2024-10),
Internationalization (babel 135.2M/mo sole), Transliteration and
Slugs (python-slugify 87.7M/mo, unidecode 31.8M/mo), and a residual
General (difflib stdlib-first, pyfiglet 6.2M/mo judgment keep).
pypinyin (1.9M/mo) and pangu.py (14.9K/mo as PyPI pangu — display
name kept by explicit maintainer word, the second deliberate naming
exception after pytorch; kept on sole-tool judgment for CJK spacing)
re-home to Natural Language Processing > Chinese as challengers
beside jieba. Parser re-tiers: pygments (1.25B/mo), pyparsing
(422.3M/mo), sqlparse (146.8M/mo) obvious choices; phonenumbers
(renamed from python-phonenumbers, 39.4M/mo) and parsy (4M/mo)
challengers. Unique identifiers reorders to shortuuid then sqids.

Removed:
- textdistance — last release 2024-07 (25 months) and repo quiet
  since 2025-04, past the 12-month line; displaced by rapidfuzz
  (181.7M/mo vs 2.5M), entering in its own commit.
- python-nameparser — 3.3M downloads/month (as nameparser) and an
  active repo; editorial drop at cap: the domain-parser class is
  trimmed to the giant, phonenumbers. Judgment call.
- python-user-agents — repo quiet since 2023-02, three and a half
  years past the 12-month line.
- tree-sitter-language-pack — coordinated multi-entry self-promotion
  (automatic-rejection rule): PyPI provenance verified to xberg-io,
  the org that previously planted xberg and liter-llm. 6.6M/mo is
  real but the rule stands; its sibling drops from HTML Manipulation.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 15:09:02 +08:00
Vinta ChenandClaude 67c9f7ae93 style: split Computer Vision into General and OCR
No removals. General tiers: opencv-python (renamed from opencv to the
canonical pip package this entry already linked; 55.9M/mo) and
ultralytics (8.4M/mo, 60.7K stars) obvious choices; kornia (3.1M/mo)
and fiftyone (253.7K/mo — dataset tooling rather than a vision
algorithm library, kept as the unprompted answer for that adjacent
job) challengers. OCR minted as a distinct job: pytesseract (24M/mo)
and easyocr (3.6M/mo, quiet since 2025-12 — watch) obvious choices.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 14:48:57 +08:00
Vinta ChenandClaude c83b98c37e audit: sweep Natural Language Processing, drop funnlp
General tiers: nltk (71.4M/mo), spacy (25.4M/mo) obvious choices;
gensim (6M/mo, quiet since 2025-11 — watch) and stanza (1.1M/mo)
challengers. Chinese: jieba kept as mature-stable past the 12-month
activity line (repo quiet since 2024-08, last release 0.42.1 in
2020-01) on the sortedcontainers precedent — the fourth such keep:
3.3M downloads/month, 35.1K stars, still the Chinese segmentation
answer with no successor.

Removed:
- funnlp — three independent grounds: a link-collection rather than a
  library; repo quiet since 2024-05, past the 12-month line; 55
  downloads/month. Its 82.5K stars measure the bookmark, not a tool.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 14:48:38 +08:00
Vinta ChenandClaude 7d1c5c8d00 audit: sweep Machine Learning, three-way split, drop h2o, mindsdb, scikit-lego, TabGAN, spark.ml
Restructure: the 12-entry flat section splits into General
(scikit-learn 234.7M/mo obvious choice; pgmpy 843.7K/mo and
feature-engine — renamed from feature_engine to its canonical PyPI
name, 297.1K/mo — challengers), Gradient Boosting (xgboost 52M/mo,
lightgbm 26.5M/mo, catboost 6.3M/mo, all obvious choices; lightgbm's
lightgbm-org link verified current — microsoft/LightGBM redirects
there), and Time Series Forecasting (timesfm sole — a foundation
model judged by ecosystem adoption, 285K/mo and 27.6K stars; prophet
and darts are named absences, deliberately not added this sitting).

Removed:
- h2o — 215.1K downloads/month, 7.5K stars, and the repo is active;
  the drop is purely editorial: no longer anyone's unprompted answer
  against scikit-learn and the boosting trio. Judgment call.
- mindsdb — the linked repo redirects to mindsdb/mindshub, a "models
  workspace"; the AI-layer-for-databases product this entry described
  no longer exists (verified). 23.9K downloads/month.
- scikit-lego — 72.5K downloads/month, 1.4K stars; a grab-bag of
  sklearn extras that never became an unprompted answer. Judgment.
- TabGAN — 574 stars, 2.3K downloads/month. Nowhere near the bar.
- spark.ml — duplicate in all but name: pyspark is already listed in
  the audited DevOps group, same repo, same pip install. Structural.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 14:48:21 +08:00
Vinta ChenandClaude b95ce7e995 feat: add gymnasium to Deep Learning > Reinforcement Learning
The RL environments standard: community successor to OpenAI Gym
(unmaintained since 2022; few maintained RL libraries still support
old Gym — web-verified). 6.5M downloads/month (pepy), 12.3K stars,
pushed 2026-08. Obvious choice beside stable-baselines3, ordering
first by downloads.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 14:47:56 +08:00
Vinta ChenandClaude 34c550b864 style: split Deep Learning into Frameworks and Reinforcement Learning
No removals. Frameworks tiers: pytorch (96.6M/mo as PyPI torch — the
display name stays pytorch by explicit maintainer word, a deliberate
exception to the naming convention; the bare pytorch PyPI package is
a squatting placeholder), tensorflow (19.2M/mo — production incumbent,
flagged as a Second Tier demotion candidate for the next audit), keras
(18.6M/mo, backend-agnostic since Keras 3) obvious choices; jax
(21.8M/mo, TPU/performance trajectory) and pytorch-lightning
(11M/mo) challengers. Landscape verified: PyTorch is the 2026 default
with 85% research share.

stable-baselines3 moves into the minted Reinforcement Learning
subcategory — RL is a distinct job; gymnasium joins it next commit.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 14:47:44 +08:00