In this categoryLocal AI · 46
- How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops
- How to Run TeleOCR Locally: One Model for Scanned and Photographed Documents
- How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks
- How to Run IndicF5 Locally: Voice Cloning for 11 Indian Languages
- How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control
- How to Run ZDTaichu5.0-9B Locally: A 9.79B Vision-Language Model Built for Spatial Reasoning
- How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description
- How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks
- MiniCPM5-2B: Install and Run OpenBMB's 2B Model That Outscores 4B Rivals
- BERT Base Uncased: Specs, How to Run It, and Why It's Still Trending in 2026
- How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing
- Stable Diffusion 3.5 Medium: What's Different From SDXL and Flux
- Qwen-Image-2512: The Text-to-Image Model That Renders Real Text
- How to Run Irodori-TTS Anime Locally: Japanese Voice Cloning and Voice Design
- How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction
- How to Run Kokoro-82M Locally: the Fastest Open-Weight TTS Model
- How to Install Ollama on macOSStart
- How to Install Ollama on Windows
- How to Run Llama 3 Locally with Ollama
- How to Pick the Right Local AI Model for Your Hardware
- Best GGUF Models to Run by VRAM Tier (8GB, 12GB, 16GB, 24GB, 48GB)
- Run LLMs in Your Browser With WebGPU: No Install, No Server (WebLLM)
- How to Use Ollama as a Drop-In OpenAI API
- GGUF vs MLX vs NVFP4: Local AI Quantization Formats Explained
- Best GPU for Running AI Locally in 2026Start
- How Much RAM Do You Need for Local AI?
- Mac vs PC for Local AI: Which Should You Choose?
- How to Build a Local AI Workstation on Any Budget
- How to Set Up Local AI on 8GB of VRAM
- How to Set Up Local AI on 12GB of VRAM
- Best Local LLM for 16 GB VRAM: Setup, Quantization and Real Speed
- How to Set Up Local AI on 24GB of VRAM
- Best Local LLM on Mac M4 16 GB: Setup, MLX vs GGUF and Real Speed
- Best Local LLM on Mac M4 Pro 48GB: Setup, Quantization and Real Speed
How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control
Fish Audio S2 Pro is a 4.56 billion parameter text-to-speech model that covers 80+ languages and lets you control prosody and emotion with plain text tags embedded in the input. Here is what the model costs in VRAM and how to install it.
Fish Audio S2 Pro is a 4.56 billion parameter open-weight text-to-speech model released March 9, 2026 by Fish Audio. It speaks 80+ languages and gives you fine-grained control over prosody and emotion by embedding natural-language tags directly in the text you want synthesized. The model ships with weights, fine-tuning code, and an SGLang-based streaming inference engine.
Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and hardware requirements against the primary source before recommending an install path.
A top-20 entrant on Hugging Face's trending list
S2 Pro entered the top 20 of Hugging Face's text-to-speech trending list. At time of writing it has 1,379 likes and 47,024 downloads on the fishaudio/s2-pro model page, plus 32,863 stars on the fishaudio/fish-speech GitHub repo. There is also a community demo Space, artificialguybr/fish-s2-pro-zero (184 likes), if you want to hear a sample before pulling the weights.
Specs at a glance
| Field | Value |
|---|---|
| Publisher | Fish Audio |
| Parameters | 4.56 billion (per the HF API's safetensors metadata) |
| Architecture | Dual-AR: 4B slow transformer + 400M fast residual transformer |
| Audio codec | RVQ, 10 codebooks, ~21 Hz frame rate |
| Training data | 10M+ hours across 80+ languages |
| Pipeline | Text to speech with optional inline emotion/prosody tags |
| Languages | 80+ (Tier 1: Japanese, English, Chinese; Tier 2: Korean, Spanish, Portuguese, Arabic, Russian, French, German) |
| License | Fish Audio Research License (non-commercial free; commercial use requires a separate license) |
| Release date | March 9, 2026 |
| GitHub stars | 32,863 (fishaudio/fish-speech) |
Hardware: what 4.56B parameters actually costs you
Fish Audio's own install docs list 24GB of GPU memory as the requirement for inference, and that is the number to plan around rather than a bf16 weight estimate. The weights alone total roughly 9.1GB at bfloat16, but the streaming pipeline, KV cache, and the 400M fast-AR pass at every timestep add enough overhead that Fish Audio does not recommend going below 24GB. That puts a 3090, 4090, or an A5000-class card in the realistic range for local use. CPU inference is possible in principle but the 4B slow-AR component makes it impractically slow for anything beyond a test clip.
| GPU VRAM | Verdict |
|---|---|
| 8GB | Not supported; well under the stated requirement |
| 12GB | Not supported; well under the stated requirement |
| 16GB | Below Fish Audio's stated requirement; expect to hit memory limits |
| 24GB | Meets Fish Audio's published inference requirement |
The production benchmark on a single H200 shows a real-time factor of 0.195, time-to-first-audio under 100ms, and throughput above 3,000 acoustic tokens per second. Those numbers will not transfer to a consumer card, but they give you a sense of how much headroom the architecture has when hardware is not the constraint.
Architecture: why it is faster than it looks
S2 Pro uses a Dual-AR design split into two components that run at different speeds. The slow AR is a 4B-parameter transformer that works along the time axis and predicts the primary semantic codebook. The fast AR is a 400M-parameter model that fills in the remaining 9 residual codebooks at each time step, recovering fine acoustic detail without running the full 4B model for every token.
Because this architecture is structurally isomorphic to a standard autoregressive LLM, it inherits all the serving optimizations that SGLang supports out of the box: continuous batching, paged KV cache, CUDA graph replay, and RadixAttention prefix caching. That is where the low RTF comes from.
Inline emotion and prosody control
S2 Pro does not use a fixed list of speaker IDs or a set of predefined style presets. Instead you embed natural-language instructions inside the text using bracket tags. The model was trained on 15,000+ unique tags, so you are not limited to a closed vocabulary.
Common examples from the model card: [whisper in small voice], [professional broadcast tone], [pitch up], [pause], [emphasis], [laughing], [excited], [angry], [sad], [interrupting]. You can place them at the word level, not just at the start of a sentence, which lets you change the delivery mid-utterance rather than applying a single style to the whole clip.
Install it
Clone the fish-speech repo, install the CUDA dependencies, and bring up the web UI or use the CLI. The project uses a Docker Compose file for the full stack.
FAQ
What is Fish Audio S2 Pro?
A 4.56 billion parameter open-weight text-to-speech model from Fish Audio that supports 80+ languages and lets you control prosody and emotion at the word level using plain-text tags embedded in your input.
How many parameters does it have?
4.56 billion, per the safetensors metadata in the Hugging Face API. The model splits into a 4B slow-AR component and a 400M fast-AR component.
What GPU do I need?
Fish Audio's install docs state 24GB of GPU memory for inference. That covers a 3090, 4090, or similar. Cards below 24GB are not the supported configuration and are likely to hit memory limits during streaming.
Which languages does it support?
80+ languages. Tier 1 (best quality) is Japanese, English, and Chinese. Tier 2 includes Korean, Spanish, Portuguese, Arabic, Russian, French, and German. Dozens of additional languages are supported at varying quality levels.
Can I use it commercially?
Not for free. The Fish Audio Research License permits research and non-commercial use at no cost. Commercial use requires a separate license from Fish Audio; contact business@fish.audio.
How fast is it on consumer hardware?
Fish Audio's published benchmark on an H200 shows an RTF of 0.195 and time-to-first-audio under 100ms. Consumer GPUs will be slower, but the Dual-AR architecture and SGLang backend are designed to keep latency low relative to model size.
Where to go from here
- Need a fully open TTS model for commercial work? Run Kokoro-82M locally covers a smaller, Apache-licensed alternative.
- Want Indian-language voice cloning from a tiny checkpoint? Run IndicF5 locally covers a 350M-parameter model that handles 11 Indian languages.
- Choosing a GPU for a local-AI box? Local AI on 24GB of VRAM covers the hardware tier S2 Pro's own docs call for.
- Compare S2 Pro against other open TTS models: browse the local AI model directory to see specs side by side.
Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)
Watch related tutorials
9:42
10:30
11:05
12:20
14:15
9:50Weekly local AI drops
New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.