audit: sweep Text Processing, dissolve General, drop textdistance, nameparser, user-agents, tree-sitter-language-pack

Restructure: the 10-entry General grab-bag dissolves — Encoding and
Unicode (chardet 224.1M/mo obvious choice, joined by
charset-normalizer next commit; ftfy 14.4M/mo kept as the fifth
mature-stable past-line keep, repo and release both 2024-10),
Internationalization (babel 135.2M/mo sole), Transliteration and
Slugs (python-slugify 87.7M/mo, unidecode 31.8M/mo), and a residual
General (difflib stdlib-first, pyfiglet 6.2M/mo judgment keep).
pypinyin (1.9M/mo) and pangu.py (14.9K/mo as PyPI pangu — display
name kept by explicit maintainer word, the second deliberate naming
exception after pytorch; kept on sole-tool judgment for CJK spacing)
re-home to Natural Language Processing > Chinese as challengers
beside jieba. Parser re-tiers: pygments (1.25B/mo), pyparsing
(422.3M/mo), sqlparse (146.8M/mo) obvious choices; phonenumbers
(renamed from python-phonenumbers, 39.4M/mo) and parsy (4M/mo)
challengers. Unique identifiers reorders to shortuuid then sqids.

Removed:
- textdistance — last release 2024-07 (25 months) and repo quiet
  since 2025-04, past the 12-month line; displaced by rapidfuzz
  (181.7M/mo vs 2.5M), entering in its own commit.
- python-nameparser — 3.3M downloads/month (as nameparser) and an
  active repo; editorial drop at cap: the domain-parser class is
  trimmed to the giant, phonenumbers. Judgment call.
- python-user-agents — repo quiet since 2023-02, three and a half
  years past the 12-month line.
- tree-sitter-language-pack — coordinated multi-entry self-promotion
  (automatic-rejection rule): PyPI provenance verified to xberg-io,
  the org that previously planted xberg and liter-llm. 6.6M/mo is
  real but the rule stands; its sibling drops from HTML Manipulation.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Vinta Chen
2026-08-16 15:09:02 +08:00
co-authored by Claude
parent 67c9f7ae93
commit da4383a06f
+15 -16
View File
@@ -213,6 +213,8 @@ _Libraries for working with human languages._
- [stanza](https://github.com/stanfordnlp/stanza) - The Stanford NLP Group's official Python library, supporting 60+ languages.
- Chinese
- [jieba](https://github.com/fxsjy/jieba) - The most popular Chinese text segmentation library.
- [pypinyin](https://github.com/mozillazg/python-pinyin) - Convert Chinese hanzi (漢字) to pinyin (拼音).
- [pangu.py](https://github.com/vinta/pangu.py) - Paranoid text spacing.
### Computer Vision
@@ -851,29 +853,26 @@ _Libraries for working with graphical user interface applications._
_Libraries for parsing and manipulating plain texts._
- General
- [babel](https://github.com/python-babel/babel) - An internationalization library for Python.
- Encoding and Unicode
- [chardet](https://github.com/chardet/chardet) - Python character encoding detector.
- [difflib](https://docs.python.org/3/library/difflib.html) - (Python standard library) Helpers for computing deltas.
- [ftfy](https://github.com/rspeer/python-ftfy) - Makes Unicode text less broken and more consistent automagically.
- [pangu.py](https://github.com/vinta/pangu.py) - Paranoid text spacing.
- General
- [difflib](https://docs.python.org/3/library/difflib.html) - (Python standard library) Helpers for computing deltas.
- [pyfiglet](https://github.com/pwaller/pyfiglet) - An implementation of figlet written in Python.
- [pypinyin](https://github.com/mozillazg/python-pinyin) - Convert Chinese hanzi (漢字) to pinyin (拼音).
- [python-slugify](https://github.com/un33k/python-slugify) - A Python slugify library that translates unicode to ASCII.
- [textdistance](https://github.com/life4/textdistance) - Compute distance between sequences with 30+ algorithms.
- [unidecode](https://github.com/avian2/unidecode) - ASCII transliterations of Unicode text.
- Unique identifiers
- [sqids](https://github.com/sqids/sqids-python) - A library for generating short unique IDs from numbers.
- [shortuuid](https://github.com/skorokithakis/shortuuid) - A generator library for concise, unambiguous and URL-safe UUIDs.
- Internationalization
- [babel](https://github.com/python-babel/babel) - An internationalization library for Python.
- Parser
- [parsy](https://github.com/python-parsy/parsy) - Easy, generic parser combinator library for creating parsers.
- [pygments](https://github.com/pygments/pygments) - A generic syntax highlighter.
- [pyparsing](https://github.com/pyparsing/pyparsing) - A general purpose framework for generating parsers.
- [python-nameparser](https://github.com/derek73/python-nameparser) - Parsing human names into their individual components.
- [python-phonenumbers](https://github.com/daviddrysdale/python-phonenumbers) - Parsing, formatting, storing and validating international phone numbers.
- [python-user-agents](https://github.com/selwin/python-user-agents) - Browser user agent parser.
- [sqlparse](https://github.com/andialbrecht/sqlparse) - A non-validating SQL parser.
- [tree-sitter-language-pack](https://github.com/xberg-io/tree-sitter-language-pack) - A comprehensive collection of tree-sitter parsers for 300+ languages, distributed as prebuilt wheels.
- [phonenumbers](https://github.com/daviddrysdale/python-phonenumbers) - Parsing, formatting, storing and validating international phone numbers.
- [parsy](https://github.com/python-parsy/parsy) - Easy, generic parser combinator library for creating parsers.
- Transliteration and Slugs
- [python-slugify](https://github.com/un33k/python-slugify) - A Python slugify library that translates unicode to ASCII.
- [unidecode](https://github.com/avian2/unidecode) - ASCII transliterations of Unicode text.
- Unique identifiers
- [shortuuid](https://github.com/skorokithakis/shortuuid) - A generator library for concise, unambiguous and URL-safe UUIDs.
- [sqids](https://github.com/sqids/sqids-python) - A library for generating short unique IDs from numbers.
### HTML Manipulation