mirror of
https://github.com/vinta/awesome-python.git
synced 2026-10-06 17:05:16 +08:00
audit: sweep HTML Manipulation, drop html-to-markdown, pyquery, tinycss2
Tiers: beautifulsoup4 (renamed from beautifulsoup — the bare PyPI name is the abandoned bs3 shim; 451.4M/mo, docs link per the PyQt precedent), lxml (401.3M/mo), xmltodict (124.5M/mo) obvious choices; markupsafe (820.5M/mo — the section's biggest raw count, but jinja-transitive infrastructure, so challenger on judgment; watch: quiet since 2025-09) and justhtml (67.8K/mo, 1.1K stars in two years — trajectory judgment on a young pure-Python HTML5 parser) challengers. Removed: - html-to-markdown — coordinated multi-entry self-promotion (automatic-rejection rule): PyPI provenance verified to xberg-io, the org's fourth planted entry overall. 1.5M downloads/month is real but the rule stands. - pyquery — 2.2M downloads/month and an active repo (pushed 2026-07); editorial drop at cap: the jQuery-style API is the least-reached-for of the keeps. Judgment call. - tinycss2 — 110.5M downloads/month is transitive (weasyprint declares it a hard dependency, verified in PyPI metadata) against 190 stars; a CSS parser mis-homed in an HTML/XML section with no better home. Judgment call. Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -881,14 +881,11 @@ _Libraries for parsing and manipulating plain texts._
|
||||
|
||||
_Libraries for working with HTML and XML._
|
||||
|
||||
- [beautifulsoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) - Providing Pythonic idioms for iterating, searching, and modifying HTML or XML.
|
||||
- [html-to-markdown](https://github.com/xberg-io/html-to-markdown) - A fast, CommonMark-compliant HTML to Markdown converter with a Rust core, tolerant of malformed HTML.
|
||||
- [justhtml](https://github.com/EmilStenstrom/justhtml/) - A pure Python HTML5 parser that just works.
|
||||
- [beautifulsoup4](https://www.crummy.com/software/BeautifulSoup/bs4/doc/) - Providing Pythonic idioms for iterating, searching, and modifying HTML or XML.
|
||||
- [lxml](https://github.com/lxml/lxml) - A very fast, easy-to-use and versatile library for handling HTML and XML.
|
||||
- [markupsafe](https://github.com/pallets/markupsafe) - Implements a XML/HTML/XHTML Markup safe string for Python.
|
||||
- [pyquery](https://github.com/gawel/pyquery) - A jQuery-like library for parsing HTML.
|
||||
- [tinycss2](https://github.com/Kozea/tinycss2) - A low-level CSS parser and generator written in Python.
|
||||
- [xmltodict](https://github.com/martinblech/xmltodict) - Working with XML feel like you are working with JSON.
|
||||
- [markupsafe](https://github.com/pallets/markupsafe) - Implements a XML/HTML/XHTML Markup safe string for Python.
|
||||
- [justhtml](https://github.com/EmilStenstrom/justhtml/) - A pure Python HTML5 parser that just works.
|
||||
|
||||
### File Format Processing
|
||||
|
||||
|
||||
Reference in New Issue
Block a user