mirror of
https://github.com/vinta/awesome-python.git
synced 2026-10-08 00:53:22 +08:00
audit: sweep Web Scraping, replace html2text with markdownify
Removed: html2text (12,324,912 downloads/month, pypistats 2026-10-02): last release 2025-04-15, last commit 2025-10-28, GPL-3.0, and markdownify does the same HTML-to-Markdown job at 52,097,079/month (about 38M after markitdown's dependency pull). Content Extraction now reads markdownify, feedparser, trafilatura. Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -396,8 +396,8 @@ _Libraries to automate web scraping and extract web content._
|
||||
- [stagehand](https://github.com/browserbase/stagehand) - A fast and token-efficient browser automation SDK to extract data and perform self-healing actions on web pages.
|
||||
- [jev-ultrafast](https://github.com/browser-use/jev-ultrafast) - A fast browser agent that picks actions from an indexed table of page elements, using a small LLM only to type text.
|
||||
- Content Extraction
|
||||
- [markdownify](https://github.com/matthewwithanm/python-markdownify) - Convert HTML to Markdown, with customizable tag handling.
|
||||
- [feedparser](https://github.com/kurtmckee/feedparser) - Universal feed parser.
|
||||
- [html2text](https://github.com/Alir3z4/html2text) - Convert HTML to Markdown-formatted text.
|
||||
- [trafilatura](https://github.com/adbar/trafilatura) - A tool for gathering text and metadata from the web, with built-in content filtering.
|
||||
|
||||
### Email
|
||||
|
||||
Reference in New Issue
Block a user