In this categoryLocal AI · 46
More guides
Local AIIntermediate

How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control

Fish Audio S2 Pro is a 4.56 billion parameter text-to-speech model that covers 80+ languages and lets you control prosody and emotion with plain text tags embedded in the input. Here is what the model costs in VRAM and how to install it.

8 minIntermediate

Fish Audio S2 Pro is a 4.56 billion parameter open-weight text-to-speech model released March 9, 2026 by Fish Audio. It speaks 80+ languages and gives you fine-grained control over prosody and emotion by embedding natural-language tags directly in the text you want synthesized. The model ships with weights, fine-tuning code, and an SGLang-based streaming inference engine.

Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and hardware requirements against the primary source before recommending an install path.

S2 Pro entered the top 20 of Hugging Face's text-to-speech trending list. At time of writing it has 1,379 likes and 47,024 downloads on the fishaudio/s2-pro model page, plus 32,863 stars on the fishaudio/fish-speech GitHub repo. There is also a community demo Space, artificialguybr/fish-s2-pro-zero (184 likes), if you want to hear a sample before pulling the weights.

Specs at a glance

FieldValue
PublisherFish Audio
Parameters4.56 billion (per the HF API's safetensors metadata)
ArchitectureDual-AR: 4B slow transformer + 400M fast residual transformer
Audio codecRVQ, 10 codebooks, ~21 Hz frame rate
Training data10M+ hours across 80+ languages
PipelineText to speech with optional inline emotion/prosody tags
Languages80+ (Tier 1: Japanese, English, Chinese; Tier 2: Korean, Spanish, Portuguese, Arabic, Russian, French, German)
LicenseFish Audio Research License (non-commercial free; commercial use requires a separate license)
Release dateMarch 9, 2026
GitHub stars32,863 (fishaudio/fish-speech)
Non-commercial by default
S2 Pro is gated on Hugging Face. You must agree to the Fish Audio Research License before downloading. Research and personal use are free; commercial use requires a separate license from Fish Audio (business@fish.audio).

Hardware: what 4.56B parameters actually costs you

Fish Audio's own install docs list 24GB of GPU memory as the requirement for inference, and that is the number to plan around rather than a bf16 weight estimate. The weights alone total roughly 9.1GB at bfloat16, but the streaming pipeline, KV cache, and the 400M fast-AR pass at every timestep add enough overhead that Fish Audio does not recommend going below 24GB. That puts a 3090, 4090, or an A5000-class card in the realistic range for local use. CPU inference is possible in principle but the 4B slow-AR component makes it impractically slow for anything beyond a test clip.

GPU VRAMVerdict
8GBNot supported; well under the stated requirement
12GBNot supported; well under the stated requirement
16GBBelow Fish Audio's stated requirement; expect to hit memory limits
24GBMeets Fish Audio's published inference requirement

The production benchmark on a single H200 shows a real-time factor of 0.195, time-to-first-audio under 100ms, and throughput above 3,000 acoustic tokens per second. Those numbers will not transfer to a consumer card, but they give you a sense of how much headroom the architecture has when hardware is not the constraint.

Architecture: why it is faster than it looks

S2 Pro uses a Dual-AR design split into two components that run at different speeds. The slow AR is a 4B-parameter transformer that works along the time axis and predicts the primary semantic codebook. The fast AR is a 400M-parameter model that fills in the remaining 9 residual codebooks at each time step, recovering fine acoustic detail without running the full 4B model for every token.

Because this architecture is structurally isomorphic to a standard autoregressive LLM, it inherits all the serving optimizations that SGLang supports out of the box: continuous batching, paged KV cache, CUDA graph replay, and RadixAttention prefix caching. That is where the low RTF comes from.

Inline emotion and prosody control

S2 Pro does not use a fixed list of speaker IDs or a set of predefined style presets. Instead you embed natural-language instructions inside the text using bracket tags. The model was trained on 15,000+ unique tags, so you are not limited to a closed vocabulary.

Common examples from the model card: [whisper in small voice], [professional broadcast tone], [pitch up], [pause], [emphasis], [laughing], [excited], [angry], [sad], [interrupting]. You can place them at the word level, not just at the start of a sentence, which lets you change the delivery mid-utterance rather than applying a single style to the whole clip.

Install it

Clone the fish-speech repo, install the CUDA dependencies, and bring up the web UI or use the CLI. The project uses a Docker Compose file for the full stack.

bash - install fish-speech
$git clone https://github.com/fishaudio/fish-speech.git
$cd fish-speech
$pip install -e .[cu129] # use .[cpu] for CPU-only
$docker compose --profile webui up # WebUI at localhost:7860
$
Verify your CUDA version first
The cu129 extra targets CUDA 12.9. If your driver targets a different CUDA version, pick the matching extra from the repo's pyproject.toml, or the install will pull incompatible binaries.

FAQ

What is Fish Audio S2 Pro?

A 4.56 billion parameter open-weight text-to-speech model from Fish Audio that supports 80+ languages and lets you control prosody and emotion at the word level using plain-text tags embedded in your input.

How many parameters does it have?

4.56 billion, per the safetensors metadata in the Hugging Face API. The model splits into a 4B slow-AR component and a 400M fast-AR component.

What GPU do I need?

Fish Audio's install docs state 24GB of GPU memory for inference. That covers a 3090, 4090, or similar. Cards below 24GB are not the supported configuration and are likely to hit memory limits during streaming.

Which languages does it support?

80+ languages. Tier 1 (best quality) is Japanese, English, and Chinese. Tier 2 includes Korean, Spanish, Portuguese, Arabic, Russian, French, and German. Dozens of additional languages are supported at varying quality levels.

Can I use it commercially?

Not for free. The Fish Audio Research License permits research and non-commercial use at no cost. Commercial use requires a separate license from Fish Audio; contact business@fish.audio.

How fast is it on consumer hardware?

Fish Audio's published benchmark on an H200 shows an RTF of 0.195 and time-to-first-audio under 100ms. Consumer GPUs will be slower, but the Dual-AR architecture and SGLang backend are designed to keep latency low relative to model size.

Where to go from here

Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)

Watch related tutorials

Free weekly email

Weekly local AI drops

New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.

Tags
#fish audio s2 pro#fishaudio s2 pro#fish speech#run fish audio locally#multilingual tts model#local text to speech