From 12f7b2b5f6fc034d3de772e95f2a7f58887b785c Mon Sep 17 00:00:00 2001 From: Vinta Chen Date: Fri, 2 Oct 2026 12:39:14 +0800 Subject: [PATCH] audit: sweep Web Scraping, replace html2text with markdownify Removed: html2text (12,324,912 downloads/month, pypistats 2026-10-02): last release 2025-04-15, last commit 2025-10-28, GPL-3.0, and markdownify does the same HTML-to-Markdown job at 52,097,079/month (about 38M after markitdown's dependency pull). Content Extraction now reads markdownify, feedparser, trafilatura. Co-Authored-By: Claude --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index edad28bb..91ced2f0 100644 --- a/README.md +++ b/README.md @@ -396,8 +396,8 @@ _Libraries to automate web scraping and extract web content._ - [stagehand](https://github.com/browserbase/stagehand) - A fast and token-efficient browser automation SDK to extract data and perform self-healing actions on web pages. - [jev-ultrafast](https://github.com/browser-use/jev-ultrafast) - A fast browser agent that picks actions from an indexed table of page elements, using a small LLM only to type text. - Content Extraction + - [markdownify](https://github.com/matthewwithanm/python-markdownify) - Convert HTML to Markdown, with customizable tag handling. - [feedparser](https://github.com/kurtmckee/feedparser) - Universal feed parser. - - [html2text](https://github.com/Alir3z4/html2text) - Convert HTML to Markdown-formatted text. - [trafilatura](https://github.com/adbar/trafilatura) - A tool for gathering text and metadata from the web, with built-in content filtering. ### Email