In this categoryLocal AI · 48
More guides
Local AIIntermediate

How to Run Freya-TTS Locally: A 183M-Parameter Turkish Text-to-Speech Model

Freya-TTS (FreyaTTS-small) is a 183 million parameter open-weight Turkish text-to-speech model built on a tokenizer-free flow-matching transformer. It needs under 2GB of VRAM and runs with a real-time factor around 0.10 on a consumer GPU. Here is what it costs to run and how to install it.

7 minIntermediate

Freya-TTS, published on Hugging Face as freyavoice/Freya-TTS and also called FreyaTTS-small, is a 183 million parameter open-weight text-to-speech model built specifically for Turkish. It is tokenizer-free at the character level, generates audio with a non-autoregressive flow-matching transformer, and runs comfortably on a single consumer GPU with about 1.5GB of VRAM. The model, inference code, and training pipeline are all released under Apache-2.0.

Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and hardware requirements against the primary source before recommending an install path.

Freya-TTS entered the top 20 of Hugging Face's text-to-speech trending list. At time of writing it has 43 likes and 5,146 downloads on the freyavoice/Freya-TTS model page, plus 160 stars on the freyavoiceai/FreyaTTS GitHub repo. The top linked community Space is EmreAkgul/turkish-tts-arena (6 likes), where you can compare it against other Turkish TTS models before installing anything.

Specs at a glance

FieldValue
PublisherFreya (freyavoiceai)
Parameters183.2 million (per the model card; 0.18B in the HF API's safetensors metadata)
ArchitectureConditional flow-matching diffusion transformer, non-autoregressive, 32-step Euler ODE
TokenizationCharacter-level, 92-symbol Turkish vocabulary, no phonemizer or G2P step
Latent spaceFrozen AudioVAE2, 64-dim at 25 Hz, 16 kHz encode / 48 kHz decode
Output48 kHz mono audio
LanguageTurkish only
LicenseApache-2.0
Release dateJuly 7, 2026
GitHub stars160 (freyavoiceai/FreyaTTS)
One name, two models
Freya's team also runs FreyaTTS-large, a higher-quality production model that powers their voice agents, but it is not released as open weights; access is commercial only, through dev@freyavoice.ai. The model id freyavoice/freya-tts you can download today is FreyaTTS-small, the one covered in this guide.

Hardware: light enough for most GPUs

At 183 million parameters, Freya-TTS asks for very little memory. On an RTX 4090 the model reports a real-time factor of 0.10 to 0.11, a time-to-first-audio of about 0.5 seconds, and 1.5GB of VRAM used, with throughput of 9.4 audio-seconds generated per second at a concurrency of 4. For comparison, the model card puts this at roughly 3.2x faster real-time factor and 3.7x less VRAM than VoxCPM2, the 2 billion parameter model whose latent space Freya-TTS reuses.

GPU VRAMVerdict
4GBComfortable fit; well above the 1.5GB measured on an RTX 4090
6GBComfortable fit, plenty of headroom for batching
8GB or moreNo constraint from this model at any reasonable batch size

You are not limited to a GPU, either. On an Apple M3 laptop CPU the model card reports a real-time factor of 0.70 in fp32, and about 0.12 end to end when routed through Core ML on Apple silicon, which is faster than the GPU figure above because it measures the full pipeline rather than just the diffusion steps.

Why it is fast: a frozen latent space and no autoregression

Freya-TTS does not predict audio sample by sample or token by token. It runs a conditional flow-matching diffusion transformer that denoises a latent representation over 32 Euler ODE steps, with no classifier-free guidance pass to double the compute. That latent representation itself is not learned by Freya-TTS at all; it comes from AudioVAE2, a frozen pretrained codec borrowed from openbmb's VoxCPM2, which compresses audio into a 64-dimensional space at 25Hz before Freya-TTS ever touches it. Keeping the codec frozen and non-autoregressive is most of why a 183M model can beat a 2B model on both speed and memory.

On the input side, the model skips phonemization entirely. It reads raw Turkish characters through a 92-symbol vocabulary, so there is no separate grapheme-to-phoneme step to install, configure, or have fail silently on an unusual word.

Accuracy: where it ranks among open Turkish TTS models

On the model's own Freya-TR-Eval benchmark, Freya-TTS scores 8.0% word error rate and 3.0% character error rate, placing 3rd of 7 open sub-1B-parameter Turkish TTS models tested. It beats XTTS-v2, which scores 11.1% WER, and F5-TTS, which scores 24.3% WER, on the same eval set.

ModelWER
Freya-TTS8.0%
XTTS-v211.1%
F5-TTS24.3%
Single speaker, no cloning
Freya-TTS ships one target voice baked into the weights. There is no speaker embedding or reference-audio input, so if you need voice cloning or multiple selectable voices, this is not the model for that job.

Install it

Clone the FreyaTTS repo and install its dependencies, then load the model straight from Hugging Face through the freyatts Python package.

bash - install freyatts
$git clone https://github.com/freyavoiceai/FreyaTTS.git
$cd FreyaTTS
$pip install -r requirements.txt
$
synthesize.py
from freyatts import FreyaTTS

tts = FreyaTTS.from_pretrained("freyavoice/freya-tts", device="cuda")
wav = tts.synthesize("Merhaba, size nasıl yardımcı olabilirim?")   # np.float32, 48 kHz
tts.save_wav(wav, "output.wav")
No GPU? Drop the device argument
The CPU and Core ML benchmarks above show this model is usable without CUDA. Swap device="cuda" for device="cpu" if you are testing on a laptop before committing to a GPU box.

FAQ

What is Freya-TTS?

A 183 million parameter open-weight text-to-speech model built for Turkish, released by Freya under Apache-2.0. It uses a tokenizer-free, character-level flow-matching transformer and reuses a frozen pretrained audio codec rather than learning its own.

How much VRAM does it need?

About 1.5GB on an RTX 4090, per the model card's own benchmark. It also runs on CPU, with a reported real-time factor of 0.70 in fp32 on an Apple M3 laptop.

Does it support languages other than Turkish?

No. Freya-TTS is trained and evaluated on Turkish only, with a 92-symbol character vocabulary specific to the language.

Can I clone a custom voice with it?

No. The open-weight model generates a single fixed voice with no speaker embedding or reference-audio input. Freya's commercial FreyaTTS-large model, available through dev@freyavoice.ai, is the one used in their production voice agents.

How does it compare to XTTS-v2 or F5-TTS?

On the model's Freya-TR-Eval benchmark, Freya-TTS scores 8.0% word error rate against 11.1% for XTTS-v2 and 24.3% for F5-TTS, ranking 3rd of 7 open sub-1B Turkish TTS models tested.

Where to go from here

Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)

Watch related tutorials

Free weekly email

Weekly local AI drops

New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.

Tags
#freya tts#freyavoice freya-tts#turkish text to speech#run freya tts locally#flow matching tts model#turkish tts model