Enterprise - Fish Audio
https://fish.audio/enterprise/ • 205 KB fetched
Open original page
Enterprise - Fish Audio
Fish Audio
Products
Solutions
Resources
Developers
Enterprise
Pricing
Contact Sales
Log in Sign Up Free
Voice infrastructure for enterprises
The expressive, controllable real-time voice model behind HeyGen, Retell, Sanas, and the next generation of voice AI builders. Production-grade across avatar video, voice agents, character apps, audio content, multilingual support, and voice-preserving translation.
Talk to sales Hear the model See pricing
Try it Code
English
Chris
S2 Pro running live. Pick a voice, type a line, hear it back. The same model behind production teams with no signup, no sales call, no demo environment.
80 +
Languages
2M +
Voice library
$15 /1M chars
Flat API rate
<150 ms
First audio (cloud)
Trusted by teams building voice in production
Six reasons voice teams switch.
Most TTS sounds fine in a demo. Fish is built for what comes after — production traffic, edge-case pronunciation, multilingual code-switching, sovereign deployments, and the kind of total cost that lets you scale instead of just survive.
01 Production-tested across enterprise scale. 02 The most expressive voice AI. Direct it, don't just pick it. 03 80+ languages, with native Asian-language depth. 04 2 million voices, one of the largest ecosystems in voice AI. 05 Self-host the model, on your infrastructure. 06 Built for how you actually grow.
Production
Listed on Artificial Analysis · public methodology
Benchmarks
Powering HeyGen, Retell, Sanas, FinalRound
Pronunciation
Custom dictionaries · numbers, names, domain terms
S2 Pro is listed on the Artificial Analysis voice leaderboard and powers production deployments at HeyGen, Retell, and Sanas — handling real traffic, edge-case pronunciation, and the kind of multi-region load that surfaces what benchmarks miss.
Production
Listed on Artificial Analysis · public methodology
Pronunciation
Custom dictionaries · numbers, names, domain terms
Benchmarks
Powering HeyGen, Retell, Sanas, FinalRound
S2 Pro is listed on the Artificial Analysis voice leaderboard and powers production deployments at HeyGen, Retell, and Sanas — handling real traffic, edge-case pronunciation, and the kind of multi-region load that surfaces what benchmarks miss.
15,000+ natural-language direction tags. Describe what you want — {warm, conversational, slight Boston accent, ending on a soft falling tone} — and Fish renders it. S2 Pro passes the Audio Turing Test with a published 0.515 score: listeners cannot reliably distinguish it from human speech. Methodology and raw audio are public.
Native-quality Mandarin, Japanese, Korean, and Cantonese — with instant code-switching across English, Mandarin, Japanese, Spanish, and Arabic. The APAC coverage other voice vendors are still promising for next quarter ships in production today.
Browse 2M+ creator-trained voices ready to use today — or clone your own from 30 seconds of audio. No slot quotas, no per-voice fees. Voice cloning with consent verification built into the workflow.
For regulated workloads, sovereign deployments, and teams that need full control of the model running in production — Fish offers self-hosting as a premium enterprise tier. Run on your VPC, your air-gapped environment, or your data center. The architecture procurement asks for and rarely gets.
$15 per million characters — flat, predictable, the same per-character rate from your first API call to your billionth. Volume discounts compound as you scale, across multiple tiers, all negotiated with one team. No seat fees. No surprise gatekeeping for production rates.
S2 Pro is listed on the Artificial Analysis voice leaderboard and powers production deployments at HeyGen, Retell, and Sanas — handling real traffic, edge-case pronunciation, and the kind of multi-region load that surfaces what benchmarks miss.
15,000+ natural-language direction tags. Describe what you want — {warm, conversational, slight Boston accent, ending on a soft falling tone} — and Fish renders it. S2 Pro passes the Audio Turing Test with a published 0.515 score: listeners cannot reliably distinguish it from human speech. Methodology and raw audio are public.
Native-quality Mandarin, Japanese, Korean, and Cantonese — with instant code-switching across English, Mandarin, Japanese, Spanish, and Arabic. The APAC coverage other voice vendors are still promising for next quarter ships in production today.
Browse 2M+ creator-trained voices ready to use today — or clone your own from 30 seconds of audio. No slot quotas, no per-voice fees. Voice cloning with consent verification built into the workflow.
For regulated workloads, sovereign deployments, and teams that need full control of the model running in production — Fish offers self-hosting as a premium enterprise tier. Run on your VPC, your air-gapped environment, or your data center. The architecture procurement asks for and rarely gets.
$15 per million characters — flat, predictable, the same per-character rate from your first API call to your billionth. Volume discounts compound as you scale, across multiple tiers, all negotiated with one team. No seat fees. No surprise gatekeeping for production rates.
Production results, not demo wins.
The headline isn't quality. It's what teams achieved after they switched. Each story is a quantified outcome, written by the customer.
Selected 3-to-1 over alternatives for voice cloning with non-American English accents.
Read the case study
Powering character-level expressiveness for Japanese AI characters inside Picto VOICE.
Read the case study
Real-time Voice Agent TTS for 10M+ users - naturalness, emotion, latency, multilingual.
Read the case study
Production voice agents with real-time orchestration for enterprise conversations.
Live interview coaching at real-time latency.
Six categories of voice product,
shipping in production today.
From avatar video to multilingual customer support - every category below is a real enterprise deployment running on Fish, not a roadmap promise.
Voice for AI agent
Character & companion apps.
Avatar & Video Voiceover
Multilingual customer support.
Mandarin · Japanese · Korean · Cantonese
Voice cloning at scale.
2M-voice ecosystem · 30-sec clone
Audio translation & dubbing.
Across all 80+ languages · code-switching
Plugs into the voice-agent stack you already use.
Drop-in support for the orchestration, telephony, and infrastructure tools voice teams ship with today. SDKs for every major language. WebSocket streaming, REST, and inbound webhook patterns documented.
Real-time pipelines
WebRTC infrastructure
Workflow automation
Voice agent platform
Telephony · SIP · SMS
Voice agent orchestration
Real-time pipelines
WebRTC infrastructure
Workflow automation
Voice agent platform
Telephony · SIP · SMS
Voice agent orchestration
The boring things that matter on a customer call.
Start at the Enterprise tier for production deployments. Volume discounts apply at higher commitment levels - talk to sales for the pricing that matches your traffic profile. For sovereign deployments, the premium self-host tier is available with a separate setup and commitment structure.
Up to 99 %
UPTIME SLA
Available on premium enterprise tier
< 150 ms
FIRST AUDIO (CLOUD)
Verified across US, EU, APAC regions
Custom
CONCURRENT STREAMS
50+ at High Volume · custom at Enterprise tier
80 +
LANGUAGES
With native-quality voices and code-switching
Built for how you actually grow.
One enterprise tier. Flat per-character pricing. Volume discounts that compound across multiple tiers as you scale - negotiated with one team, in one contract.
Start at the Enterprise tier for production deployments. Volume discounts apply at higher commitment levels - talk to sales for the pricing that matches your traffic profile. For sovereign deployments, the premium self-host tier is available with a separate setup and commitment structure.
Plan Inclusions
Enterprise Plan
Terms & Notes
Starting Price
From $999 / month
Volume discounts at higher commitment tiers
TTS · S2 Pro
$15 / 1M characters
Billed in UTF-8 bytes · about 180K English words per 1M
TTS · S1
$15 / 1M characters
Same flat rate as S2 Pro
ASR · transcribe-1
$0.36 / audio hour
Duration rounded up to the nearest second
Concurrency
Custom
50+ at High Volume tier · custom at Enterprise
Voices
Unlimited
No slot quotas · no per-voice fees
SLA
Up to 99%
Available on premium enterprise tier
Support
Dedicated Slack channel
Compliance SOC2 / HIPAA upon requests
Self-host premium
From $10K setup + $10K / month
12-month commit · VPC · on-prem · air-gapped · sovereign cloud
Volume discounts available across multiple tiers - contact sales for pricing that matches your traffic profile. Public pricing reflects Enterprise tier entry. Larger commitments unlock further discounts on a per-customer basis.
Ready when you are.
Talk to our team about your deployment. We'll come prepared.
Talk to sales
Frequently asked questions
Where is my data stored? Do you support U.S., EU, and APAC residency? By default your data stays in the United States, hosted on Google Cloud with Cloudflare R2 storage, and inference runs from edge regions in the U.S. and Asia-Pacific (Tokyo) so your users get low latency wherever they are. For compliance-bound workloads, enterprise contracts can switch on Zero Data Retention, which means request text and audio are never written to disk. And if your data has to stay inside a specific country or region, the self-hosted enterprise tier runs fully inside your own infrastructure, so nothing ever leaves your environment.
Can you support large-scale deployments and traffic spikes? Yes, and at serious volume. Capacity is provisioned as concurrent generations that scale with your contract, and we already have production customers running more than 1,000 concurrent generations. A Rust edge gateway serves inference across multiple GPU regions, so when your traffic surges our team can lift your limits the same day. You scale up without ever queuing behind a support ticket.
What security certifications do you have? Security runs through every layer of the platform. Our SOC 2 Type II audit is currently underway, and the report will be available to customers under NDA once it is complete. Zero Data Retention is available on enterprise contracts, so request payloads are never persisted, and the self-hosted tier keeps every byte of your data inside your own environment. We also support HIPAA-aligned configurations and can sign a BAA for qualifying healthcare workloads, and independent penetration testing runs as part of our ongoing compliance program.
Do you offer engineering support for custom deployments? Absolutely. Enterprise customers get a direct line to our engineering team, not a ticketing queue, on whatever channel suits how your team works. We ship integration-specific features and protocol extensions for individual customers on a regular basis, and we stand up self-hosted deployments with you end to end, from first setup through go-live.
Do you support SSO and RBAC? Yes, with fine-grained control from day one. Role-based access control lets you assign owner, admin, and member roles at the team level, plus manager, contributor, and viewer roles at the workspace level, so everyone has exactly the access they should. Single sign-on works today through Google and GitHub OAuth.
Can we fine-tune models on our data, or use our own voices? Both, and on your terms. You can spin up private voice clones from as little as 10 seconds of reference audio, 30 seconds or more for the best results, instantly through the API or the web UI, and they stay fully private to your team. For deeper engagements, we also fine-tune custom models on your own data.
What about migration from another voice vendor? Migrating to Fish Audio is straightforward, and most teams are surprised how quickly it goes. Your existing voices come across by recreating them from reference audio, our Python, TypeScript, and Go SDKs and WebSocket streaming API cover the integration patterns you already rely on, and our engineering team runs the cutover alongside you so production never skips a beat.
Products
* Text-to-Speech
* Speech-to-Text
* Voice Cloning
* Voice Changer
* Story Studio
* Audio Separation
* Audio Translation
* Sound Effects
* AI Voice Generator
Solutions
* For Startups
* For Students
* Audiobooks
* Voiceovers
* Character Voices
* Conversational Chatbots
Research
* OpenAudio
* Fish Audio S2
* Fish Audio S1
* Fish Speech
* Fish Diffusion
Resources
* Discovery
* Guide
* API Reference
* Voice Library
* Compare Us
* Affiliate
* Pricing
Company
* GitHub
* Blog
* Support
Products
* Text-to-Speech
* Speech-to-Text
* Voice Cloning
* Voice Changer
* Story Studio
* Audio Separation
* Audio Translation
* Sound Effects
* AI Voice Generator
Solutions
* For Startups
* For Students
* Audiobooks
* Voiceovers
* Character Voices
* Conversational Chatbots
Research
* OpenAudio
* Fish Audio S2
* Fish Audio S1
* Fish Speech
* Fish Diffusion
Resources
* Discovery
* Guide
* API Reference
* Voice Library
* Compare Us
* Affiliate
* Pricing
Company
* GitHub
* Blog
* Support
Fish Audio © 2026 Hanabi AI Inc. All rights reserved. Privacy Policy Terms of Service Report Abuse
English
Links found on this page
- Fish Audio [direct]
- Products [direct]
- Resources [direct]
- Enterprise [direct]
- Pricing [direct]
- Contact Sales [direct]
- Hear the model [direct]
- See pricing [direct]
- Read the case study [direct]
- Read the case study [direct]
- Read the case study [direct]
- Real-time pipelines [direct]
- WebRTC infrastructure [direct]
- Workflow automation [direct]
- Voice agent platform [direct]
- Telephony · SIP · SMS [direct]
- Voice agent orchestration [direct]
- Speech-to-Text [direct]
- Voice Cloning [direct]
- Voice Changer [direct]
- Story Studio [direct]
- Audio Separation [direct]
- Audio Translation [direct]
- Sound Effects [direct]
- AI Voice Generator [direct]
- For Startups [direct]
- For Students [direct]
- Audiobooks [direct]
- Voiceovers [direct]
- Character Voices [direct]
- Conversational Chatbots [direct]
- OpenAudio [direct]
- Fish Audio S2 [direct]
- Fish Audio S1 [direct]
- Fish Speech [direct]
- Fish Diffusion [direct]
- Discovery [direct]
- Guide [direct]
- API Reference [direct]
- Voice Library [direct]
- Compare Us [direct]
- Affiliate [direct]
- GitHub [direct]
- Blog [direct]
- Privacy Policy [direct]
- Terms of Service [direct]