audit: sweep HTML Manipulation, drop html-to-markdown, pyquery, tinycss2

Tiers: beautifulsoup4 (renamed from beautifulsoup — the bare PyPI
name is the abandoned bs3 shim; 451.4M/mo, docs link per the PyQt
precedent), lxml (401.3M/mo), xmltodict (124.5M/mo) obvious choices;
markupsafe (820.5M/mo — the section's biggest raw count, but
jinja-transitive infrastructure, so challenger on judgment; watch:
quiet since 2025-09) and justhtml (67.8K/mo, 1.1K stars in two
years — trajectory judgment on a young pure-Python HTML5 parser)
challengers.

Removed:
- html-to-markdown — coordinated multi-entry self-promotion
  (automatic-rejection rule): PyPI provenance verified to xberg-io,
  the org's fourth planted entry overall. 1.5M downloads/month is
  real but the rule stands.
- pyquery — 2.2M downloads/month and an active repo (pushed
  2026-07); editorial drop at cap: the jQuery-style API is the
  least-reached-for of the keeps. Judgment call.
- tinycss2 — 110.5M downloads/month is transitive (weasyprint
  declares it a hard dependency, verified in PyPI metadata) against
  190 stars; a CSS parser mis-homed in an HTML/XML section with no
  better home. Judgment call.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Vinta Chen
2026-08-16 15:09:52 +08:00
co-authored by Claude
parent db9c262342
commit ad698c2876
+3 -6
View File
@@ -881,14 +881,11 @@ _Libraries for parsing and manipulating plain texts._
_Libraries for working with HTML and XML._
- [beautifulsoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) - Providing Pythonic idioms for iterating, searching, and modifying HTML or XML.
- [html-to-markdown](https://github.com/xberg-io/html-to-markdown) - A fast, CommonMark-compliant HTML to Markdown converter with a Rust core, tolerant of malformed HTML.
- [justhtml](https://github.com/EmilStenstrom/justhtml/) - A pure Python HTML5 parser that just works.
- [beautifulsoup4](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) - Providing Pythonic idioms for iterating, searching, and modifying HTML or XML.
- [lxml](https://github.com/lxml/lxml) - A very fast, easy-to-use and versatile library for handling HTML and XML.
- [markupsafe](https://github.com/pallets/markupsafe) - Implements a XML/HTML/XHTML Markup safe string for Python.
- [pyquery](https://github.com/gawel/pyquery) - A jQuery-like library for parsing HTML.
- [tinycss2](https://github.com/Kozea/tinycss2) - A low-level CSS parser and generator written in Python.
- [xmltodict](https://github.com/martinblech/xmltodict) - Working with XML feel like you are working with JSON.
- [markupsafe](https://github.com/pallets/markupsafe) - Implements a XML/HTML/XHTML Markup safe string for Python.
- [justhtml](https://github.com/EmilStenstrom/justhtml/) - A pure Python HTML5 parser that just works.
### File Format Processing