Published: April 26, 2025
24
184
1.7k

🚨 Two undergrads with almost no AI background just built a speech model called "Dia" that rivals Google’s NotebookLM and ElevenLabs. And it's open-source. Here's how they pulled it off and why it matters:

Image in tweet by Brendan Jowett

The market for synthetic speech is booming. Startups raised $398M for voice AI last year. ElevenLabs leads the pack but challengers like PlayAI, Sesame, and now, Nari Labs, are gaining ground fast.

Image in tweet by Brendan Jowett

Nari Labs was founded by two students in Korea. They started learning about speech AI just three months ago. Inspired by NotebookLM, they wanted a model with more control over tone, style, and nonverbal cues. (Dia from Nari Labs Vs ElevenLabs - listen to this 🤯)

They built Dia: a 1.6-billion-parameter model trained using Google’s TPU Research Cloud. Dia can: - Generate podcast-style conversations - Clone voices - Insert laughs, coughs, disfluencies into dialogue

Image in tweet by Brendan Jowett

You can run Dia on any modern PC with 10GB VRAM. It’s hosted on Hugging Face and GitHub, free to use. By default, it generates random voices unless you prompt it with a style or upload a voice to clone.

Image in tweet by Brendan Jowett

What’s next for Nari? They’re planning a full synthetic voice platform with a social layer on top of Dia. They’re also working on multilingual support—and bigger, more powerful models.

Two undergrads. Three months. One open-source voice model that could disrupt the entire synthetic speech market. The AI wave is getting harder and harder to predict.

P.S. AI Voice Agents don’t just answer calls, they close deals, set appointments, and follow up instantly. Want to automate your sales pipeline? Let’s talk. Learn more here: https://www.inflate.agency/?ut...

Share this thread

Read on Twitter

View original thread

Navigate thread

1/8