In this categoryLocal AI · 48
- How to Run Nemotron 3 Diarization Locally: Streaming Speaker Tagging in 99M Parameters
- How to Run Freya-TTS Locally: A 183M-Parameter Turkish Text-to-Speech Model
- How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops
- How to Run TeleOCR Locally: One Model for Scanned and Photographed Documents
- How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks
- How to Run IndicF5 Locally: Voice Cloning for 11 Indian Languages
- How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control
- How to Run ZDTaichu5.0-9B Locally: A 9.79B Vision-Language Model Built for Spatial Reasoning
- How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description
- How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks
- MiniCPM5-2B: Install and Run OpenBMB's 2B Model That Outscores 4B Rivals
- BERT Base Uncased: Specs, How to Run It, and Why It's Still Trending in 2026
- How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing
- Stable Diffusion 3.5 Medium: What's Different From SDXL and Flux
- Qwen-Image-2512: The Text-to-Image Model That Renders Real Text
- How to Run Irodori-TTS Anime Locally: Japanese Voice Cloning and Voice Design
- How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction
- How to Run Kokoro-82M Locally: the Fastest Open-Weight TTS Model
- How to Install Ollama on macOSStart
- How to Install Ollama on Windows
- How to Run Llama 3 Locally with Ollama
- How to Pick the Right Local AI Model for Your Hardware
- Best GGUF Models to Run by VRAM Tier (8GB, 12GB, 16GB, 24GB, 48GB)
- Run LLMs in Your Browser With WebGPU: No Install, No Server (WebLLM)
- How to Use Ollama as a Drop-In OpenAI API
- GGUF vs MLX vs NVFP4: Local AI Quantization Formats Explained
- Best GPU for Running AI Locally in 2026Start
- How Much RAM Do You Need for Local AI?
- Mac vs PC for Local AI: Which Should You Choose?
- How to Build a Local AI Workstation on Any Budget
- How to Set Up Local AI on 8GB of VRAM
- How to Set Up Local AI on 12GB of VRAM
- Best Local LLM for 16 GB VRAM: Setup, Quantization and Real Speed
- How to Set Up Local AI on 24GB of VRAM
- Best Local LLM on Mac M4 16 GB: Setup, MLX vs GGUF and Real Speed
- Best Local LLM on Mac M4 Pro 48GB: Setup, Quantization and Real Speed
How to Run Freya-TTS Locally: A 183M-Parameter Turkish Text-to-Speech Model
Freya-TTS (FreyaTTS-small) is a 183 million parameter open-weight Turkish text-to-speech model built on a tokenizer-free flow-matching transformer. It needs under 2GB of VRAM and runs with a real-time factor around 0.10 on a consumer GPU. Here is what it costs to run and how to install it.
Freya-TTS, published on Hugging Face as freyavoice/Freya-TTS and also called FreyaTTS-small, is a 183 million parameter open-weight text-to-speech model built specifically for Turkish. It is tokenizer-free at the character level, generates audio with a non-autoregressive flow-matching transformer, and runs comfortably on a single consumer GPU with about 1.5GB of VRAM. The model, inference code, and training pipeline are all released under Apache-2.0.
Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and hardware requirements against the primary source before recommending an install path.
A new top-20 entrant on Hugging Face's trending list
Freya-TTS entered the top 20 of Hugging Face's text-to-speech trending list. At time of writing it has 43 likes and 5,146 downloads on the freyavoice/Freya-TTS model page, plus 160 stars on the freyavoiceai/FreyaTTS GitHub repo. The top linked community Space is EmreAkgul/turkish-tts-arena (6 likes), where you can compare it against other Turkish TTS models before installing anything.
Specs at a glance
| Field | Value |
|---|---|
| Publisher | Freya (freyavoiceai) |
| Parameters | 183.2 million (per the model card; 0.18B in the HF API's safetensors metadata) |
| Architecture | Conditional flow-matching diffusion transformer, non-autoregressive, 32-step Euler ODE |
| Tokenization | Character-level, 92-symbol Turkish vocabulary, no phonemizer or G2P step |
| Latent space | Frozen AudioVAE2, 64-dim at 25 Hz, 16 kHz encode / 48 kHz decode |
| Output | 48 kHz mono audio |
| Language | Turkish only |
| License | Apache-2.0 |
| Release date | July 7, 2026 |
| GitHub stars | 160 (freyavoiceai/FreyaTTS) |
Hardware: light enough for most GPUs
At 183 million parameters, Freya-TTS asks for very little memory. On an RTX 4090 the model reports a real-time factor of 0.10 to 0.11, a time-to-first-audio of about 0.5 seconds, and 1.5GB of VRAM used, with throughput of 9.4 audio-seconds generated per second at a concurrency of 4. For comparison, the model card puts this at roughly 3.2x faster real-time factor and 3.7x less VRAM than VoxCPM2, the 2 billion parameter model whose latent space Freya-TTS reuses.
| GPU VRAM | Verdict |
|---|---|
| 4GB | Comfortable fit; well above the 1.5GB measured on an RTX 4090 |
| 6GB | Comfortable fit, plenty of headroom for batching |
| 8GB or more | No constraint from this model at any reasonable batch size |
You are not limited to a GPU, either. On an Apple M3 laptop CPU the model card reports a real-time factor of 0.70 in fp32, and about 0.12 end to end when routed through Core ML on Apple silicon, which is faster than the GPU figure above because it measures the full pipeline rather than just the diffusion steps.
Why it is fast: a frozen latent space and no autoregression
Freya-TTS does not predict audio sample by sample or token by token. It runs a conditional flow-matching diffusion transformer that denoises a latent representation over 32 Euler ODE steps, with no classifier-free guidance pass to double the compute. That latent representation itself is not learned by Freya-TTS at all; it comes from AudioVAE2, a frozen pretrained codec borrowed from openbmb's VoxCPM2, which compresses audio into a 64-dimensional space at 25Hz before Freya-TTS ever touches it. Keeping the codec frozen and non-autoregressive is most of why a 183M model can beat a 2B model on both speed and memory.
On the input side, the model skips phonemization entirely. It reads raw Turkish characters through a 92-symbol vocabulary, so there is no separate grapheme-to-phoneme step to install, configure, or have fail silently on an unusual word.
Accuracy: where it ranks among open Turkish TTS models
On the model's own Freya-TR-Eval benchmark, Freya-TTS scores 8.0% word error rate and 3.0% character error rate, placing 3rd of 7 open sub-1B-parameter Turkish TTS models tested. It beats XTTS-v2, which scores 11.1% WER, and F5-TTS, which scores 24.3% WER, on the same eval set.
| Model | WER |
|---|---|
| Freya-TTS | 8.0% |
| XTTS-v2 | 11.1% |
| F5-TTS | 24.3% |
Install it
Clone the FreyaTTS repo and install its dependencies, then load the model straight from Hugging Face through the freyatts Python package.
from freyatts import FreyaTTS
tts = FreyaTTS.from_pretrained("freyavoice/freya-tts", device="cuda")
wav = tts.synthesize("Merhaba, size nasıl yardımcı olabilirim?") # np.float32, 48 kHz
tts.save_wav(wav, "output.wav")FAQ
What is Freya-TTS?
A 183 million parameter open-weight text-to-speech model built for Turkish, released by Freya under Apache-2.0. It uses a tokenizer-free, character-level flow-matching transformer and reuses a frozen pretrained audio codec rather than learning its own.
How much VRAM does it need?
About 1.5GB on an RTX 4090, per the model card's own benchmark. It also runs on CPU, with a reported real-time factor of 0.70 in fp32 on an Apple M3 laptop.
Does it support languages other than Turkish?
No. Freya-TTS is trained and evaluated on Turkish only, with a 92-symbol character vocabulary specific to the language.
Can I clone a custom voice with it?
No. The open-weight model generates a single fixed voice with no speaker embedding or reference-audio input. Freya's commercial FreyaTTS-large model, available through dev@freyavoice.ai, is the one used in their production voice agents.
How does it compare to XTTS-v2 or F5-TTS?
On the model's Freya-TR-Eval benchmark, Freya-TTS scores 8.0% word error rate against 11.1% for XTTS-v2 and 24.3% for F5-TTS, ranking 3rd of 7 open sub-1B Turkish TTS models tested.
Where to go from here
- Need multilingual TTS instead of a Turkish-only model? Run Fish Audio S2 Pro locally covers an 80+ language model with inline emotion control.
- Want a fully open, Apache-licensed TTS model that is not language-locked? Run Kokoro-82M locally covers a similarly compact alternative.
- Pairing TTS with speaker tagging for a full voice pipeline? Run Nemotron 3 Diarization locally covers a 99M-parameter diarization model that needs no GPU.
- Compare Freya-TTS against other open TTS models: browse the local AI model directory to see specs side by side.
Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)
Watch related tutorials
9:42
10:30
11:05
9:50
13:30
16:45Weekly local AI drops
New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.