Fish Audio
The most expressive, emotionally controllable real-time voice model
About
5 seconds of audio to clone a voice is either a superpower or a subpoena waiting to happen. Curious what the consent guardrails actually look like in the API.
One thing that'd close the loop: a voice library marketplace where creators license their cloned voice with revenue share. Fish becomes the platform instead of just the model.
The tweet buries the lede by putting the raise before the product. Would've hooked harder leading with the 5-second clone demo and letting the funding be the mic drop.
Word-level emotion tags sound fun until I'm debugging why 'excited' sounds like 'mildly concerned'. What's the rate limit on streaming and do you expose SSML-ish controls or a custom schema?
Every voice startup this quarter has claimed to be Nx faster and 1/Yth the cost of Eleven Labs. At some point the benchmark becomes a personality trait.
Hot take: expressive TTS is a feature, not a company. The moment the frontier labs bundle this into their audio APIs, half the category evaporates.
Would love to see the actual eval methodology behind 'most expressive'. MOS scores? Preference tests? Or is it vibes on a Tuesday afternoon?
The launch video pacing is genuinely tight, whoever cut those emotion demos deserves a raise. Not sure the 'happy vs sad' A/B needed to loop three times though.
ok wait, if it's 2x faster than Cartesia in real-time streaming, that changes what you can build for live agents. This is the number I actually care about.
Naive question but if I clone my own voice and later delete the sample, does the model still remember me? Asking because the docs weren't clear.