IA + agentes

6 min · Jul 25

Fish.audio: the open TTS and voice cloning platform behind modern voice agents

What Fish.audio is, how its TTS and voice cloning work, and what open voice models mean for a business that wants a natural-sounding voice agent.

If you are exploring how to give an AI agent a voice, sooner or later you will run into fish.audio. It is one of the fastest-growing text-to-speech (TTS) platforms among developers, largely because it leans on open models, voice cloning, and pricing that is more accessible than closed premium options. This guide covers what it is, how it works, and what it means if you want your business to answer calls in a natural voice.

What is Fish.audio?

Fish.audio (https://fish.audio) is a speech synthesis platform: you give it text and it returns spoken audio. What sets it apart is not only the website for generating audio, but its family of open models — known as Fish-Speech and, more recently, the OpenAudio line — built for developers to use via API or even run on their own infrastructure. Around that sits a community voice marketplace, where voices created by other users are shared and discovered.

How TTS and voice cloning work

The basic TTS flow is easy to grasp: the model turns a string of text into an audio waveform that sounds like a person speaking. The interesting part of Fish.audio is voice cloning — from a relatively short audio sample, the system can generate a new voice that mimics the timbre and style of that sample.

Text-to-speech: a model turns plain text into a spoken audio waveform.
  • Text-to-speech (TTS): you write the script and the model reads it in a synthetic voice.
  • Voice cloning: you upload a sample and get a custom voice with that timbre.
  • Voice marketplace: you browse and reuse voices created by the community.
  • API access: you plug audio generation into your own product or workflow.
  • Self-hosting: because the models are open, you can run them on your own infrastructure if you need to.

Cloning a voice carries responsibility. Only use voices you have permission or rights to, and be transparent with your contacts when they are talking to an AI agent. Naturalness should never be used to mislead people about who — or what — is on the other end.

Open models vs. closed premium voices

This is where it pays to be honest about the trade-offs. Closed, premium voice options tend to lead on naturalness and consistency, especially for specific languages and accents. Open models like those from Fish.audio have narrowed that gap considerably and bring clear advantages: lower cost, more control, and the option to self-host.

Comparing voices is not just "which sounds better" — weigh naturalness, consistency, language, cost, and control.
  • In favor of open: lower cost, transparency, the ability to self-host, and not depending on a single vendor.
  • In favor of closed premium: usually wins on naturalness and fine consistency, particularly in certain languages and accents.
  • The language factor: results in a given language or accent can vary a lot between models — test with your own script before deciding.

Where does it fit in a real voice agent?

Generating a nice audio clip is one thing; holding a real phone call is another. A voice agent that answers the phone needs far more than good TTS: it has to listen (speech recognition), understand and improvise (the language model), reply with low latency, handle interruptions when the person talks over it, and know when to hand the call to a human.

  1. Speech-to-text (STT) to transcribe what the person says.
  2. A language model that understands, reasons, and decides what to say.
  3. TTS (such as Fish.audio) to turn that reply into a natural voice.
  4. Low latency and interruption handling so the conversation flows.
  5. Escalation to a human when the situation calls for it.

TTS — whether Fish.audio or another — is just one piece of that puzzle. Assembling the whole system yourself is a serious engineering project: you have to orchestrate voice, latency, conversation memory, and integrations with your business.

From a standalone voice to a complete agent

If your end goal is for your business to answer calls and messages in a natural voice, Fish.audio is worth knowing — but you do not have to become an integrator to get there. Blind Agents brings the voice, the language, and the integrations together into one agent that answers WhatsApp, phone calls, and email, with no code, and escalates to someone on your team when needed. You focus on the business; we focus on making it sound right.

Fish.audio: the open TTS and voice cloning platform behind modern voice agents · Blind Agents