audit: sweep Audio & Video Processing, re-home TTS to Speech, drop birdnet, matchering

Restructure: gtts and KittenTTS re-home to AI and Agents > Speech — TTS
belongs with the speech models, not audio processing. gtts (3M/mo,
a 2014 service wrapper, quiet since 2026-04) lands as Second Tier;
kittentts (renamed to its canonical PyPI name, 15.3K stars within a
year of creation, pip count 3.9K/mo under-measures model adoption)
as challenger. Speech is now at exact cap (5/5) — future additions
there require displacement.

Audio keeps pydub (21.4M/mo) and librosa (11.7M/mo) as obvious
choices. Video unchanged: moviepy (7M/mo) obvious choice, vidgear
(25.3K/mo) challenger. Metadata reorders to tiers: mutagen (8.2M/mo)
obvious choice; tinytag (5.2M/mo), beets (67.4K/mo, kept as challenger
rather than minting a sole-entry subcategory — judgment call).

Removed:
- birdnet — 7.1K downloads/month, 1.7K stars; an applied bioacoustics
  tool with a genuinely small audience, not an audio-processing
  choice. Judgment call.
- matchering — 16.1K downloads/month, 2.6K stars; a single-task
  automated-mastering tool with a small audience, no trajectory.
  Judgment call.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Vinta Chen
2026-08-16 13:14:37 +08:00
co-authored by Claude
parent 9970486a2a
commit 8f9c3884bf
+4 -6
View File
@@ -168,6 +168,8 @@ _Libraries for building AI applications, LLM integrations, and autonomous agents
- [openai-whisper](https://github.com/openai/whisper) - A general-purpose automatic speech recognition model trained on 680k hours of multilingual and multitask supervised data.
- [funasr](https://github.com/modelscope/FunASR) - Industrial-grade speech recognition toolkit with 170x realtime speed, 50+ languages, speaker diarization, and emotion detection.
- [vibevoice](https://github.com/microsoft/VibeVoice) - A family of open-source voice AI models from Microsoft for text-to-speech and long-form speech recognition.
- [gtts](https://github.com/pndurette/gTTS) - Python library and CLI tool for converting text to speech using Google Translate TTS.
- [kittentts](https://github.com/KittenML/KittenTTS) - Lightweight ONNX text-to-speech library with small CPU-friendly models.
### Deep Learning
@@ -939,19 +941,15 @@ _Libraries for manipulating images._
_Libraries for manipulating audio, video, and their metadata._
- Audio
- [birdnet](https://github.com/birdnet-team/BirdNET-Analyzer) - Deep learning framework for acoustic species detection; identifies bird species from audio recordings using TensorFlow.
- [gtts](https://github.com/pndurette/gTTS) - Python library and CLI tool for converting text to speech using Google Translate TTS.
- [KittenTTS](https://github.com/KittenML/KittenTTS) - Lightweight ONNX text-to-speech library with small CPU-friendly models.
- [librosa](https://github.com/librosa/librosa) - Python library for audio and music analysis.
- [matchering](https://github.com/sergree/matchering) - A library for automated reference audio mastering.
- [pydub](https://github.com/jiaaro/pydub) - Manipulate audio with a simple and easy high level interface.
- [librosa](https://github.com/librosa/librosa) - Python library for audio and music analysis.
- Video
- [moviepy](https://github.com/Zulko/moviepy) - A module for script-based movie editing with many formats, including animated GIFs.
- [vidgear](https://github.com/abhiTronix/vidgear) - Most Powerful multi-threaded Video Processing framework.
- Metadata
- [beets](https://github.com/beetbox/beets) - A music library manager and [MusicBrainz](https://musicbrainz.org/) tagger.
- [mutagen](https://github.com/quodlibet/mutagen) - A Python module to handle audio metadata.
- [tinytag](https://github.com/tinytag/tinytag) - A library for reading music meta data of MP3, OGG, FLAC and Wave files.
- [beets](https://github.com/beetbox/beets) - A music library manager and [MusicBrainz](https://musicbrainz.org/) tagger.
### Game Development