From ad698c28765d7c9125a286baffe3b27d9d4bc010 Mon Sep 17 00:00:00 2001 From: Vinta Chen Date: Sun, 16 Aug 2026 15:09:52 +0800 Subject: [PATCH] audit: sweep HTML Manipulation, drop html-to-markdown, pyquery, tinycss2 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Tiers: beautifulsoup4 (renamed from beautifulsoup — the bare PyPI name is the abandoned bs3 shim; 451.4M/mo, docs link per the PyQt precedent), lxml (401.3M/mo), xmltodict (124.5M/mo) obvious choices; markupsafe (820.5M/mo — the section's biggest raw count, but jinja-transitive infrastructure, so challenger on judgment; watch: quiet since 2025-09) and justhtml (67.8K/mo, 1.1K stars in two years — trajectory judgment on a young pure-Python HTML5 parser) challengers. Removed: - html-to-markdown — coordinated multi-entry self-promotion (automatic-rejection rule): PyPI provenance verified to xberg-io, the org's fourth planted entry overall. 1.5M downloads/month is real but the rule stands. - pyquery — 2.2M downloads/month and an active repo (pushed 2026-07); editorial drop at cap: the jQuery-style API is the least-reached-for of the keeps. Judgment call. - tinycss2 — 110.5M downloads/month is transitive (weasyprint declares it a hard dependency, verified in PyPI metadata) against 190 stars; a CSS parser mis-homed in an HTML/XML section with no better home. Judgment call. Co-Authored-By: Claude --- README.md | 9 +++------ 1 file changed, 3 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index 4e54b7eb..4c8203d0 100644 --- a/README.md +++ b/README.md @@ -881,14 +881,11 @@ _Libraries for parsing and manipulating plain texts._ _Libraries for working with HTML and XML._ -- [beautifulsoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) - Providing Pythonic idioms for iterating, searching, and modifying HTML or XML. -- [html-to-markdown](https://github.com/xberg-io/html-to-markdown) - A fast, CommonMark-compliant HTML to Markdown converter with a Rust core, tolerant of malformed HTML. -- [justhtml](https://github.com/EmilStenstrom/justhtml/) - A pure Python HTML5 parser that just works. +- [beautifulsoup4](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) - Providing Pythonic idioms for iterating, searching, and modifying HTML or XML. - [lxml](https://github.com/lxml/lxml) - A very fast, easy-to-use and versatile library for handling HTML and XML. -- [markupsafe](https://github.com/pallets/markupsafe) - Implements a XML/HTML/XHTML Markup safe string for Python. -- [pyquery](https://github.com/gawel/pyquery) - A jQuery-like library for parsing HTML. -- [tinycss2](https://github.com/Kozea/tinycss2) - A low-level CSS parser and generator written in Python. - [xmltodict](https://github.com/martinblech/xmltodict) - Working with XML feel like you are working with JSON. +- [markupsafe](https://github.com/pallets/markupsafe) - Implements a XML/HTML/XHTML Markup safe string for Python. +- [justhtml](https://github.com/EmilStenstrom/justhtml/) - A pure Python HTML5 parser that just works. ### File Format Processing