TechnicalNews

FREE Unlimited AI Voices | Better Than ElevenLabs (Microsoft Banned It)

Helena Liu

The video demonstrates how to use Vibe Voice, a free open-source AI text-to-speech model that Microsoft removed from the market due to deepfake concerns, to generate realistic AI audio locally without subscription fees. The presenter shows how to install the model, create multi-speaker podcasts, and clone a personal voice for custom audio generation.

Summary

The presenter explains that Microsoft developed Vibe Voice, a highly realistic text-to-speech model, but feared releasing it due to deepfake concerns. After finally open-sourcing it over a year ago, Microsoft pulled the entire project offline within weeks due to these same concerns. However, community members had already archived copies of the GitHub repository before removal. The video provides step-by-step instructions for installing the Vibe Voice model locally on a computer using cloud code, offering both a smaller 1.5B parameter model and a larger 7B model. The presenter demonstrates the tool's capabilities, including pre-loaded voice options, multi-speaker podcast generation with up to four speakers, and multilingual audio generation combining English and Chinese. A key feature showcased is voice cloning: by providing a 30-second audio sample recorded on an iPhone, users can create a custom voice replica that the model learns and reproduces. The presenter converts their own voice memo from M4A to WAV format and successfully clones it as a usable voice option in the interface. The resulting cloned voice is compared favorably to paid services like ElevenLabs. The presenter also mentions Quinn as an alternative open-source Chinese voice generation model. Use cases discussed include custom audio for customer purchases, blog-to-audio conversion, audiobook generation, and podcast creation—essentially replacing paid services like ElevenLabs and Notebook LM at no cost. The video concludes with promotion of the presenter's free AI agent course and business consulting services.

Key Insights

  • Microsoft developed Vibe Voice but was afraid to release it initially because the generated voices were too realistic and could flood the market with deepfakes
  • Microsoft removed Vibe Voice from the internet within three weeks of release due to excessive deepfake concerns, but community members had already saved copies of the entire GitHub repository before takedown
  • The official Microsoft version of Vibe Voice was stripped of all text-to-speech functionality, leaving only transcription capabilities, so users must install a community-preserved version to generate audio
  • Vibe Voice can generate audio with proper intonation, rhythm, expression, and natural pauses between words, and supports multiple languages simultaneously in the same audio output
  • Users can clone their own voice by providing just a 30-second audio sample recorded in a quiet room, which the model then learns to reproduce with high fidelity comparable to paid services

Topics

Vibe Voice open-source AI text-to-speech modelMicrosoft's removal of Vibe Voice from marketLocal AI audio generation without subscriptionsVoice cloning and customizationMulti-speaker podcast generationMultilingual audio synthesisComparison with paid alternatives (ElevenLabs, Notebook LM)GitHub repository archiving by community membersInstallation and setup procedures

Transcript

[0:00] Forget 11 loves. Microsoft just released a free voice calling tool so good that they actually had to take it off of the market three weeks after it was released. Okay, here's the tea, guys. So, over a year ago, Microsoft developed Vibe Voice, but they were too afraid to release it because the voices that it was able to generate were too realistic and they were afraid it would flood the market with deep fakes. Now, last year they finally released the project and they open sourced the entire thing. But within weeks, there were way too much [0:31] concerns about deep fake. So, Microsoft pulled the entire project off of the internet. But here's the thing, though.…

Full transcript available for MurmurCast members

Sign Up to Access

More from Helena Liu

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.