In this categoryLocal AI · 46
More guides
Local AIIntermediate

How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops

Audio8 ASR Infinite is a 4.09 billion parameter streaming speech recognition model from Edge0 that transcribes Chinese and English audio 24/7 with a rolling KV cache, so memory and latency stay flat no matter how long the recording runs. Here is what it costs in VRAM and how to run it.

7 minIntermediate

Audio8 ASR Infinite is a 4.09 billion parameter open-weight speech recognition model from Edge0, released September 21, 2026. It transcribes Chinese and English audio as a live stream rather than in fixed chunks, and a rolling KV cache lets it run continuously without memory growing or transcription drifting, even across a 24 hour recording. The weights are Apache 2.0 licensed.

Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and VRAM requirements against the primary source before recommending an install path.

Audio8 ASR Infinite entered the top 20 of Hugging Face's overall trending list within a week of release. At time of writing it has 1,008 likes and 19,434 downloads on the Hugging Face repo, its GitHub repo has 91 stars, and the top linked demo, embedl/hfviewer, has 139 likes.

Specs at a glance

FieldValue
PublisherEdge0
Parameters4.09 billion (BF16 safetensors, per the HF API)
PipelineAutomatic speech recognition, native streaming
LanguagesChinese and English
Audio clockSelectable 80 / 120 / 160 ms
Transcription delayConfigurable, 240 to 560 ms depending on clock
ArchitectureVoxtral Realtime 4B audio tower + Qwen2.5-3B-Instruct decoder, DSM-style streaming
Weights size8.17GB (model.safetensors, BF16) plus a separate semantic VAD head file
LicenseApache 2.0
Release dateSeptember 21, 2026
GitHub stars91
License is fully open
Apache 2.0 covers the weights, so there is no gated request form or commercial tier to clear before you can use it.

What makes it different: it does not restart

Most streaming ASR models cap out at a fixed context window and either truncate or reset once audio runs past it. Audio8 ASR Infinite's checkpoint has a native context of 30 seconds, but a rolling KV cache lets it keep decoding past that limit indefinitely, holding memory and latency constant instead of letting either grow with recording length. Edge0 also built in a semantic voice activity detector with four separate time horizons (0.5, 1.0, 2.0 and 3.0 seconds) that is meant to tell a thinking pause or a stutter apart from an actual end of turn, which is where acoustic-only VAD tends to cut people off mid-sentence.

The tradeoff is a clock you set yourself. A faster 80ms clock decodes 12.5 times a second and can run at a 240ms delay, which suits live captioning or a voice agent. A slower 160ms clock decodes about 6.25 times a second at a 320 to 480ms delay, trading responsiveness for a lighter compute load. All three clocks are post-trained and documented, so you are not guessing at an unsupported configuration.

Hardware: what 4.09B parameters actually costs you

The published weights are almost entirely BF16: 4.09 billion parameters at two bytes each works out to the 8.17GB model.safetensors file Edge0 ships, plus a small additional file for the semantic VAD heads. That is a fraction of the footprint of the large mixture-of-experts language models we usually cover here, and it fits comfortably on a single consumer GPU with 12GB of VRAM or more once you leave room for the rolling KV cache and audio preprocessing overhead. This is a model you can run on a workstation card, not a cluster.

Rolling KV cache means VRAM does not creep up
Because the KV cache rolls instead of growing, a 5 minute clip and a 5 hour stream cost roughly the same VRAM once decoding is underway. Budget for the weights plus a fixed cache allowance, not for recording length.

Benchmark scores

Edge0 compares Audio8 ASR Infinite against Voxtral-Mini-4B-Realtime-2602 and nemotron-3.5-asr-streaming-0.6b at a 480ms delay on an 80ms clock. Error rates are in percent, lower is better:

Test setMetricAudio8 ASR InfiniteVoxtral-Mini-4B-Realtime-2602nemotron-3.5-asr-streaming-0.6b
aishell1/testCER1.75016.79512.927 @ 560ms
aishell4/testCER2.89316.45614.677 @ 560ms
librispeech test.cleanWER3.0422.2103.353 @ 560ms
librispeech test.otherWER6.8085.5527.140 @ 560ms
Average3.62310.253 (2 sets)9.524

The gap is widest on the Chinese-language Aishell sets, where Audio8 ASR Infinite's error rate is roughly a tenth of Voxtral's. On the English-only LibriSpeech sets it trades the lead with Voxtral, coming in slightly behind on both clean and noisy audio. Edge0 reports this as a greedy decode with no repetition loops and no dropped trailing words, which matters more for a streaming model than a marginal error-rate difference.

Run it with the torch streaming example

For a local test without standing up a server, Edge0's GitHub repo ships a torch-based simulated-streaming example that runs directly against a downloaded checkpoint:

zsh - transcribe a file with the torch streaming example
$python -m audio8_asr_infinite.examples.torch_streaming_decode \
--checkpoint /path/to/checkpoint \
--audio sample.wav --language zh --transcription-delay-ms 480
$

For always-on use, Edge0 documents Docker Compose as the canonical deployment path, since it also serves a web demo alongside the vLLM-backed realtime endpoint:

zsh - run the 24/7 vLLM deployment
$cd docker
$AUDIO8_MODEL_DIR=/path/to/checkpoint docker compose up -d
$

That brings up a web demo at http://localhost:8080/ (or https://localhost:8443/ over TLS) and a realtime websocket endpoint you can drive from the command line with the included client:

zsh - stream a file to the vLLM endpoint
$python -m audio8_asr_infinite.examples.vllm_realtime_client \
--ws-url ws://127.0.0.1:18191/v1/realtime \
--audio sample.wav --language zh --target-delay-ms 480 --pace
$
Match the delay to the clock
target_delay_ms and transcription-delay-ms must be an integer multiple of whatever audio clock you selected. The post-trained combinations are 240/320/480/560ms at an 80ms clock, 240/480ms at 120ms, and 320/480ms at 160ms. Other values will run but Edge0 has not validated accuracy outside that table.

FAQ

What is Audio8 ASR Infinite used for?

Live speech-to-text for Chinese and English audio, built for use cases that run continuously, such as always-on captioning, meeting transcription, or a voice agent's ear, rather than one-off file transcription.

How many parameters does it have?

4.09 billion, per the safetensors metadata in the Hugging Face API, made up of a Voxtral Realtime audio tower feeding a Qwen2.5-3B-Instruct decoder.

What GPU do I need to run it?

The BF16 weights are 8.17GB, so a single consumer GPU with 12GB of VRAM or more has enough headroom for the model plus its rolling KV cache and audio preprocessing. It does not need the multi-GPU setups that large language models in this size class often require.

Can it really transcribe audio 24/7 without slowing down?

That is the model's core design goal. Its native context is 30 seconds, but a rolling KV cache keeps memory and latency flat past that point instead of growing with recording length, which is what makes continuous operation practical on fixed hardware.

Is it free to use commercially?

Yes. The weights are licensed under Apache 2.0, with no separate commercial tier or approval step required.

Where to go from here

Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)

Watch related tutorials

Free weekly email

Weekly local AI drops

New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.

Tags
#audio8 asr infinite#edge0 audio8#run audio8 asr infinite locally#streaming speech recognition model#local speech to text#24/7 transcription model