Brief IA

Fish Audio raises $52M to revolutionize AI voices

💼 Business & Startups·Tom Levy·

Fish Audio raises $52M to revolutionize AI voices

Fish Audio raises $52M to revolutionize AI voices
Key Takeaways
1Fish Audio, based in Palo Alto, has raised $52 million to develop AI voice models.
2The startup has already attracted over 8 million users and generates $21 million in annual revenue.
3Fish Audio has automated the removal of unauthorized voices, reducing the turnaround time to three minutes.
💡Why it mattersThis funding round allows Fish Audio to strengthen its position against increasing competition in the AI voice technology sector.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

A Rapidly Expanding Market for AI Voice Models

The sector of AI-generated voice models is experiencing rapid growth. Creative applications require more expressive AI voices, while businesses are looking to automate functions such as customer support and sales with more controllable models.

Fish Audio: A Response to Diverse Needs

Based in Palo Alto, Fish Audio strives to meet these varied needs with an impressive library of over 15,000 natural language controls. Since its inception last year, the startup has attracted more than 8 million users for its open-source or hosted models, generating an annual recurring revenue of $21 million.

Significant Funding to Support Growth

To sustain this momentum, Fish Audio announced it has raised $52 million in a seed funding round. This round was led by Coreline Ventures and Capital Today, with participation from several other investors such as 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.

The Beginnings of Fish Audio and Its Evolution

Fish Audio started as a personal project by Shijia Liao, a former researcher at Nvidia. Frustrated by the lack of expressive synthetic voices, Liao developed a voice model on a single GPU, which he then shared as open source. The Fish Speech repository on GitHub has since garnered over 31,000 stars and is used by developers and creators from various backgrounds.

Diverse Models for Varied Needs

Over the past year, Fish Audio has launched five different models: four for speech generation and one for speech-to-text conversion. While three of these models are available as open source, the most advanced model, S2.1 Pro, is accessible only through a paid API.

Tailored Offerings for Creators and Businesses

Fish Audio offers monthly subscriptions that allow creators to access a defined number of minutes of generation and voice cloning features. For businesses, a dedicated version of its APIs is available, already adopted by organizations such as HeyGen, Sanas, and LiveKit.

Creator Trust: A Crucial Challenge

Rissa Cao, CEO and co-founder of Fish Audio, emphasizes the importance of addressing the specific needs of businesses, whether it’s realism for AI avatars or expressive voices for video games. However, creator trust is essential, especially after incidents where voices were uploaded without consent.

Fish Audio has built its voice library by inviting users to submit their own voices to train its models. Users are compensated if their voices are used, which raised concerns when some creators claimed their voices had been uploaded without their consent. To address these concerns, the startup initially implemented a DMCA takedown process, although some users found it slow.

An Automated Takedown Process

To resolve these issues, Fish Audio has established an automated process that allows creators to remove their voices in under three minutes after verification. Despite this, the risk remains that voices may be uploaded without the artists' knowledge.

Community at the Heart of the Strategy

Osuke Honda from Coreline Ventures stresses the importance of trust and transparency for a sustainable community model. He stated that a community-driven model only works when creators trust the platform, highlighting the need for clear licensing terms and revenue-sharing models.

A Vision for the Future

Although Fish Audio initially operated without external funding, the growing interest from investors has led to this fundraising to develop more advanced models and expand its customer base. The startup plans to launch an audio understanding model and a speech-to-speech conversion model soon.

A Competitive Market

The speech generation sector is highly competitive, with companies like ElevenLabs, WellSaid, Cartesia, Speechify, Async (formerly Podcastle), and Krisp. Rico Mallozzi from 359 Capital believes that Fish Audio's fine controls and cost-effective model training will give it a competitive edge.

Conclusion

With its recent funding, Fish Audio is well-positioned to continue innovating and standing out in the field of AI voice technologies while strengthening the trust of creators and businesses.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.