audit: sweep Web Scraping, replace html2text with markdownify

Removed: html2text (12,324,912 downloads/month, pypistats 2026-10-02): last release 2025-04-15, last commit 2025-10-28, GPL-3.0, and markdownify does the same HTML-to-Markdown job at 52,097,079/month (about 38M after markitdown's dependency pull). Content Extraction now reads markdownify, feedparser, trafilatura.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Vinta Chen
2026-10-02 12:39:14 +08:00
co-authored by Claude
parent 1c1994251d
commit 12f7b2b5f6
+1 -1
View File
@@ -396,8 +396,8 @@ _Libraries to automate web scraping and extract web content._
- [stagehand](https://github.com/browserbase/stagehand) - A fast and token-efficient browser automation SDK to extract data and perform self-healing actions on web pages.
- [jev-ultrafast](https://github.com/browser-use/jev-ultrafast) - A fast browser agent that picks actions from an indexed table of page elements, using a small LLM only to type text.
- Content Extraction
- [markdownify](https://github.com/matthewwithanm/python-markdownify) - Convert HTML to Markdown, with customizable tag handling.
- [feedparser](https://github.com/kurtmckee/feedparser) - Universal feed parser.
- [html2text](https://github.com/Alir3z4/html2text) - Convert HTML to Markdown-formatted text.
- [trafilatura](https://github.com/adbar/trafilatura) - A tool for gathering text and metadata from the web, with built-in content filtering.
### Email