In this categoryLocal AI · 46
- How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops
- How to Run TeleOCR Locally: One Model for Scanned and Photographed Documents
- How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks
- How to Run IndicF5 Locally: Voice Cloning for 11 Indian Languages
- How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control
- How to Run ZDTaichu5.0-9B Locally: A 9.79B Vision-Language Model Built for Spatial Reasoning
- How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description
- How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks
- MiniCPM5-2B: Install and Run OpenBMB's 2B Model That Outscores 4B Rivals
- BERT Base Uncased: Specs, How to Run It, and Why It's Still Trending in 2026
- How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing
- Stable Diffusion 3.5 Medium: What's Different From SDXL and Flux
- Qwen-Image-2512: The Text-to-Image Model That Renders Real Text
- How to Run Irodori-TTS Anime Locally: Japanese Voice Cloning and Voice Design
- How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction
- How to Run Kokoro-82M Locally: the Fastest Open-Weight TTS Model
- How to Install Ollama on macOSStart
- How to Install Ollama on Windows
- How to Run Llama 3 Locally with Ollama
- How to Pick the Right Local AI Model for Your Hardware
- Best GGUF Models to Run by VRAM Tier (8GB, 12GB, 16GB, 24GB, 48GB)
- Run LLMs in Your Browser With WebGPU: No Install, No Server (WebLLM)
- How to Use Ollama as a Drop-In OpenAI API
- GGUF vs MLX vs NVFP4: Local AI Quantization Formats Explained
- Best GPU for Running AI Locally in 2026Start
- How Much RAM Do You Need for Local AI?
- Mac vs PC for Local AI: Which Should You Choose?
- How to Build a Local AI Workstation on Any Budget
- How to Set Up Local AI on 8GB of VRAM
- How to Set Up Local AI on 12GB of VRAM
- Best Local LLM for 16 GB VRAM: Setup, Quantization and Real Speed
- How to Set Up Local AI on 24GB of VRAM
- Best Local LLM on Mac M4 16 GB: Setup, MLX vs GGUF and Real Speed
- Best Local LLM on Mac M4 Pro 48GB: Setup, Quantization and Real Speed
How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops
Audio8 ASR Infinite is a 4.09 billion parameter streaming speech recognition model from Edge0 that transcribes Chinese and English audio 24/7 with a rolling KV cache, so memory and latency stay flat no matter how long the recording runs. Here is what it costs in VRAM and how to run it.
Audio8 ASR Infinite is a 4.09 billion parameter open-weight speech recognition model from Edge0, released September 21, 2026. It transcribes Chinese and English audio as a live stream rather than in fixed chunks, and a rolling KV cache lets it run continuously without memory growing or transcription drifting, even across a 24 hour recording. The weights are Apache 2.0 licensed.
Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and VRAM requirements against the primary source before recommending an install path.
A new top-20 entrant on Hugging Face's trending list
Audio8 ASR Infinite entered the top 20 of Hugging Face's overall trending list within a week of release. At time of writing it has 1,008 likes and 19,434 downloads on the Hugging Face repo, its GitHub repo has 91 stars, and the top linked demo, embedl/hfviewer, has 139 likes.
Specs at a glance
| Field | Value |
|---|---|
| Publisher | Edge0 |
| Parameters | 4.09 billion (BF16 safetensors, per the HF API) |
| Pipeline | Automatic speech recognition, native streaming |
| Languages | Chinese and English |
| Audio clock | Selectable 80 / 120 / 160 ms |
| Transcription delay | Configurable, 240 to 560 ms depending on clock |
| Architecture | Voxtral Realtime 4B audio tower + Qwen2.5-3B-Instruct decoder, DSM-style streaming |
| Weights size | 8.17GB (model.safetensors, BF16) plus a separate semantic VAD head file |
| License | Apache 2.0 |
| Release date | September 21, 2026 |
| GitHub stars | 91 |
What makes it different: it does not restart
Most streaming ASR models cap out at a fixed context window and either truncate or reset once audio runs past it. Audio8 ASR Infinite's checkpoint has a native context of 30 seconds, but a rolling KV cache lets it keep decoding past that limit indefinitely, holding memory and latency constant instead of letting either grow with recording length. Edge0 also built in a semantic voice activity detector with four separate time horizons (0.5, 1.0, 2.0 and 3.0 seconds) that is meant to tell a thinking pause or a stutter apart from an actual end of turn, which is where acoustic-only VAD tends to cut people off mid-sentence.
The tradeoff is a clock you set yourself. A faster 80ms clock decodes 12.5 times a second and can run at a 240ms delay, which suits live captioning or a voice agent. A slower 160ms clock decodes about 6.25 times a second at a 320 to 480ms delay, trading responsiveness for a lighter compute load. All three clocks are post-trained and documented, so you are not guessing at an unsupported configuration.
Hardware: what 4.09B parameters actually costs you
The published weights are almost entirely BF16: 4.09 billion parameters at two bytes each works out to the 8.17GB model.safetensors file Edge0 ships, plus a small additional file for the semantic VAD heads. That is a fraction of the footprint of the large mixture-of-experts language models we usually cover here, and it fits comfortably on a single consumer GPU with 12GB of VRAM or more once you leave room for the rolling KV cache and audio preprocessing overhead. This is a model you can run on a workstation card, not a cluster.
Benchmark scores
Edge0 compares Audio8 ASR Infinite against Voxtral-Mini-4B-Realtime-2602 and nemotron-3.5-asr-streaming-0.6b at a 480ms delay on an 80ms clock. Error rates are in percent, lower is better:
| Test set | Metric | Audio8 ASR Infinite | Voxtral-Mini-4B-Realtime-2602 | nemotron-3.5-asr-streaming-0.6b |
|---|---|---|---|---|
| aishell1/test | CER | 1.750 | 16.795 | 12.927 @ 560ms |
| aishell4/test | CER | 2.893 | 16.456 | 14.677 @ 560ms |
| librispeech test.clean | WER | 3.042 | 2.210 | 3.353 @ 560ms |
| librispeech test.other | WER | 6.808 | 5.552 | 7.140 @ 560ms |
| Average | 3.623 | 10.253 (2 sets) | 9.524 |
The gap is widest on the Chinese-language Aishell sets, where Audio8 ASR Infinite's error rate is roughly a tenth of Voxtral's. On the English-only LibriSpeech sets it trades the lead with Voxtral, coming in slightly behind on both clean and noisy audio. Edge0 reports this as a greedy decode with no repetition loops and no dropped trailing words, which matters more for a streaming model than a marginal error-rate difference.
Run it with the torch streaming example
For a local test without standing up a server, Edge0's GitHub repo ships a torch-based simulated-streaming example that runs directly against a downloaded checkpoint:
For always-on use, Edge0 documents Docker Compose as the canonical deployment path, since it also serves a web demo alongside the vLLM-backed realtime endpoint:
That brings up a web demo at http://localhost:8080/ (or https://localhost:8443/ over TLS) and a realtime websocket endpoint you can drive from the command line with the included client:
FAQ
What is Audio8 ASR Infinite used for?
Live speech-to-text for Chinese and English audio, built for use cases that run continuously, such as always-on captioning, meeting transcription, or a voice agent's ear, rather than one-off file transcription.
How many parameters does it have?
4.09 billion, per the safetensors metadata in the Hugging Face API, made up of a Voxtral Realtime audio tower feeding a Qwen2.5-3B-Instruct decoder.
What GPU do I need to run it?
The BF16 weights are 8.17GB, so a single consumer GPU with 12GB of VRAM or more has enough headroom for the model plus its rolling KV cache and audio preprocessing. It does not need the multi-GPU setups that large language models in this size class often require.
Can it really transcribe audio 24/7 without slowing down?
That is the model's core design goal. Its native context is 30 seconds, but a rolling KV cache keeps memory and latency flat past that point instead of growing with recording length, which is what makes continuous operation practical on fixed hardware.
Is it free to use commercially?
Yes. The weights are licensed under Apache 2.0, with no separate commercial tier or approval step required.
Where to go from here
- Need a matching local text-to-speech model to pair with it? Run Fish Audio S2 Pro locally covers a similarly sized open-weight TTS model with emotion control.
- Not sure a 12GB card is enough for your whole stack? How to set up local AI on 8GB of VRAM shows what fits at an even tighter budget.
- Want to compare it against other open models first? Browse the local AI model directory to see specs side by side before you commit disk space to a download.
Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)
Watch related tutorials
9:42
10:30
11:05
12:20
14:15
9:50Weekly local AI drops
New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.