audit: sweep Natural Language Processing, drop funnlp

General tiers: nltk (71.4M/mo), spacy (25.4M/mo) obvious choices;
gensim (6M/mo, quiet since 2025-11 — watch) and stanza (1.1M/mo)
challengers. Chinese: jieba kept as mature-stable past the 12-month
activity line (repo quiet since 2024-08, last release 0.42.1 in
2020-01) on the sortedcontainers precedent — the fourth such keep:
3.3M downloads/month, 35.1K stars, still the Chinese segmentation
answer with no successor.

Removed:
- funnlp — three independent grounds: a link-collection rather than a
  library; repo quiet since 2024-05, past the 12-month line; 55
  downloads/month. Its 82.5K stars measure the bookmark, not a tool.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Vinta Chen
2026-08-16 14:48:38 +08:00
co-authored by Claude
parent 7d1c5c8d00
commit c83b98c37e
+1 -2
View File
@@ -207,12 +207,11 @@ _Libraries for Machine Learning. Also see [awesome-machine-learning](https://git
_Libraries for working with human languages._
- General
- [gensim](https://github.com/piskvorky/gensim) - Topic Modeling for Humans.
- [nltk](https://github.com/nltk/nltk) - A leading platform for building Python programs to work with human language data.
- [spacy](https://github.com/explosion/spaCy) - A library for industrial-strength natural language processing in Python and Cython.
- [gensim](https://github.com/piskvorky/gensim) - Topic Modeling for Humans.
- [stanza](https://github.com/stanfordnlp/stanza) - The Stanford NLP Group's official Python library, supporting 60+ languages.
- Chinese
- [funnlp](https://github.com/fighting41love/funNLP) - A collection of tools and datasets for Chinese NLP.
- [jieba](https://github.com/fxsjy/jieba) - The most popular Chinese text segmentation library.
### Computer Vision