September 19, 2026
Explore VoiceStudio, a local ElevenLabs alternative for voice cloning, dubbing, dictation, transcription, and audiobooks across 646 languages.
YuE2 turns lyrics and style prompts into editable musical scores, frontier-quality songs, zero-shot covers, and agent-driven revisions.
VoiceStudio brings voice cloning, dubbing, dictation, transcription, and audiobooks to your hardware with 16 TTS and 11 ASR engines.
StemDeck is an open-source, local-first stem extraction studio that isolates vocals, drums, bass, guitar, and piano with no cloud subscriptions or uploads.
Generate stereo 48 kHz music locally with HOT-Step CPP, a feature-rich C++/GGML interface for ACE-Step, audio tools, plugins, and model training.
Inflect v2 delivers complete 24 kHz English TTS in two compact models (3.96M and 9.36M params) with no external vocoder. Explore architecture, evaluation, and CPU deployment.
KittenTTS delivers high-quality text-to-speech in models as small as 25MB, running entirely on CPU. Learn how to use it, its API, and why it matters for edge AI.
CrispASR is a unified C++ speech engine built on ggml, supporting 53 ASR and 51 TTS models with zero Python dependencies. Run Cohere Transcribe, Parakeet, Voxtral, Qwen3, and more from a single CLI.
FluidAudio is a Swift SDK for running state-of-the-art speech-to-text, text-to-speech, speaker diarization, and voice activity detection entirely on-device using Apple's Neural Engine. This post explores its capabilities, architecture, and how to integrate it into your apps.
Explore SpeechRecognition, a versatile Python library supporting multiple speech recognition engines and APIs, both online and offline, with practical examples and troubleshooting tips.
Explore Speech Swift, an open-source toolkit for ASR, TTS, speech-to-speech, VAD, and diarization on Apple Silicon using MLX and CoreML.
Miso TTS 8B is a state-of-the-art, open-source text-to-speech model with 8 billion parameters, offering highly emotive speech generation and voice cloning capabilities.