From fdf66eb1f70489a131c7544d9acd37d2411bc997 Mon Sep 17 00:00:00 2001 From: Vinta Chen Date: Sun, 16 Aug 2026 01:54:20 +0800 Subject: [PATCH] audit: sweep Web Scraping to the shortlist cap First Audit of the section; maintainer adjudicated 2026-08-16. Both Subcategories reordered by downloads/month (Frameworks: browser-use, scrapy, crawl4ai; Content Extraction: feedparser, html2text with a watch - repo quiet since 2025-10 - and trafilatura). Removed (downloads are PyPI last-30-day via ClickPy, fetched 2026-08-16): - mechanicalsoup (162K/month): stateful-browsing era superseded by browser automation - crawlberg (10.6K/month): xberg-io coordinated self-promotion (the org behind the standing rejection rule); repo created 2026-03, 155 stars - website-downloader (352/month): hobby mirroring script, 174 stars - micawber (209K/month): oEmbed extraction, small audience next to the kept three - sumy (168K/month): summarization, not content extraction - wrong job for the Subcategory Co-Authored-By: Claude --- README.md | 7 +------ 1 file changed, 1 insertion(+), 6 deletions(-) diff --git a/README.md b/README.md index 1a1821ad..2e750e48 100644 --- a/README.md +++ b/README.md @@ -362,16 +362,11 @@ _Libraries to automate web scraping and extract web content._ - Frameworks - [browser-use](https://github.com/browser-use/browser-use) - Make websites accessible for AI agents with easy browser automation. - - [crawl4ai](https://github.com/unclecode/crawl4ai) - An open-source, LLM-friendly web crawler that provides lightning-fast, structured data extraction specifically designed for AI agents. - - [crawlberg](https://github.com/xberg-io/crawlberg) - A high-performance web crawling engine with a Rust core, headless-browser fallback, and built-in robots.txt and sitemap parsing. - - [mechanicalsoup](https://github.com/MechanicalSoup/MechanicalSoup) - A Python library for automating interaction with websites. - [scrapy](https://github.com/scrapy/scrapy) - A fast high-level screen scraping and web crawling framework. - - [website-downloader](https://github.com/PKHarsimran/website-downloader) - A modern wget --mirror / HTTrack alternative that turns whole websites into browsable offline copies. + - [crawl4ai](https://github.com/unclecode/crawl4ai) - An open-source, LLM-friendly web crawler that provides lightning-fast, structured data extraction specifically designed for AI agents. - Content Extraction - [feedparser](https://github.com/kurtmckee/feedparser) - Universal feed parser. - [html2text](https://github.com/Alir3z4/html2text) - Convert HTML to Markdown-formatted text. - - [micawber](https://github.com/coleifer/micawber) - A small library for extracting rich content from URLs. - - [sumy](https://github.com/miso-belica/sumy) - A module for automatic summarization of text documents and HTML pages. - [trafilatura](https://github.com/adbar/trafilatura) - A tool for gathering text and metadata from the web, with built-in content filtering. ### Email