Back to directory
AI & ML · AI voice & audio

Fish Audio

The most expressive, emotionally controllable real-time voice model

Expressive TTS and Audio Generation| Click link to claim credits 👇 https://t.co/gzh4M4dZHY
San Francisco, CA24K followers
TLVC Rating
Hook
Editing / Creativity
Copy
Sentiment of launch
Distribution strategy
Community Rating
No ratings yet
Your rating
Sign in to rate this launch.

About

Fish Audio is publicly launching S2.1 Pro, a real-time text-to-speech model aimed at developers building voice agents, dubbing pipelines, and creator tools. The model clones a voice from about five seconds of reference audio and exposes word-level controls for emotion, intonation, and pacing through natural-language bracket tags like [whisper], [excited], or [laugh], letting builders shape delivery inside the script itself rather than re-recording takes. Fish Audio S2.1 Pro is a flagship text-to-speech model built for highly expressive, low-latency speech generation. It supports natural-language bracket cues for emotion and delivery control, multi-speaker dialogue in a single generation, 80+ languages with automatic language detection, and realtime streaming with very fast time to first audio. The launch is paired with a seed round that signals how competitive the voice layer of the AI stack has become. Fish Audio, the AI voice platform for expressive real-time text-to-speech, voice cloning, and voice agents, today announced $52 million in seed funding led by Coreline Ventures and Capital Today, with participation from 359 Capital, Play Time , among others. The company was founded by CEO Rissa Cao and chief scientist Shijia Liao, who formerly worked as a video researcher at Nvidia Corp. and is a lifelong fan of Japanese anime, explained that he grew tired of having to listen to the flat and monotonously robotic synthetic voices common in earlier TTS systems. Both appear in the launch demo cloning their own voices to place a pizza order through a live agent. For operators, the pitch is cost and latency against incumbents, with Fish claiming it runs roughly twice as fast as Cartesia at about one-sixth the price of ElevenLabs. Adoption is already broad across the AI voice ecosystem, with customers including HeyGen, Retell, Sanas, LiveKit, and OpenArt, and the underlying S2 family is open-sourced on GitHub for teams that want to self-host. Fish Audio S2 Pro is a leading text-to-speech model with fine-grained inline control of prosody and emotion. Trained on over 10M+ hours of audio data across 80+ languages, it combines reinforcement learning alignment with a Dual-Autoregressive architecture, which is the technical backbone S2.1 Pro extends for production use.
Tags
1M-3MSeedCinematicB2D (developers)Product launchB2BGlobalUSVertical AIFunding announcementFounder-led
Comments (10)
Sign in to join the discussion.
Renata Okafor15d ago

5 seconds of audio to clone a voice is either a superpower or a subpoena waiting to happen. Curious what the consent guardrails actually look like in the API.

Hiroshi Tanabe15d ago

One thing that'd close the loop: a voice library marketplace where creators license their cloned voice with revenue share. Fish becomes the platform instead of just the model.

Tomek P.15d ago

The tweet buries the lede by putting the raise before the product. Would've hooked harder leading with the 5-second clone demo and letting the funding be the mic drop.

yuki15d ago

Word-level emotion tags sound fun until I'm debugging why 'excited' sounds like 'mildly concerned'. What's the rate limit on streaming and do you expose SSML-ish controls or a custom schema?

Priya Ramaswamy15d ago

Every voice startup this quarter has claimed to be Nx faster and 1/Yth the cost of Eleven Labs. At some point the benchmark becomes a personality trait.

Magnus Hedlund15d ago

Hot take: expressive TTS is a feature, not a company. The moment the frontier labs bundle this into their audio APIs, half the category evaporates.

DeShawn Whitfield15d ago

Would love to see the actual eval methodology behind 'most expressive'. MOS scores? Preference tests? Or is it vibes on a Tuesday afternoon?

Lina Sørvik15d ago

The launch video pacing is genuinely tight, whoever cut those emotion demos deserves a raise. Not sure the 'happy vs sad' A/B needed to loop three times though.

kwame15d ago

ok wait, if it's 2x faster than Cartesia in real-time streaming, that changes what you can build for live agents. This is the number I actually care about.

Anaïs T.15d ago

Naive question but if I clone my own voice and later delete the sample, does the model still remember me? Asking because the docs weren't clear.