In this categoryLocal AI · 46
- How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops
- How to Run TeleOCR Locally: One Model for Scanned and Photographed Documents
- How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks
- How to Run IndicF5 Locally: Voice Cloning for 11 Indian Languages
- How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control
- How to Run ZDTaichu5.0-9B Locally: A 9.79B Vision-Language Model Built for Spatial Reasoning
- How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description
- How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks
- MiniCPM5-2B: Install and Run OpenBMB's 2B Model That Outscores 4B Rivals
- BERT Base Uncased: Specs, How to Run It, and Why It's Still Trending in 2026
- How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing
- Stable Diffusion 3.5 Medium: What's Different From SDXL and Flux
- Qwen-Image-2512: The Text-to-Image Model That Renders Real Text
- How to Run Irodori-TTS Anime Locally: Japanese Voice Cloning and Voice Design
- How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction
- How to Run Kokoro-82M Locally: the Fastest Open-Weight TTS Model
- How to Install Ollama on macOSStart
- How to Install Ollama on Windows
- How to Run Llama 3 Locally with Ollama
- How to Pick the Right Local AI Model for Your Hardware
- Best GGUF Models to Run by VRAM Tier (8GB, 12GB, 16GB, 24GB, 48GB)
- Run LLMs in Your Browser With WebGPU: No Install, No Server (WebLLM)
- How to Use Ollama as a Drop-In OpenAI API
- GGUF vs MLX vs NVFP4: Local AI Quantization Formats Explained
- Best GPU for Running AI Locally in 2026Start
- How Much RAM Do You Need for Local AI?
- Mac vs PC for Local AI: Which Should You Choose?
- How to Build a Local AI Workstation on Any Budget
- How to Set Up Local AI on 8GB of VRAM
- How to Set Up Local AI on 12GB of VRAM
- Best Local LLM for 16 GB VRAM: Setup, Quantization and Real Speed
- How to Set Up Local AI on 24GB of VRAM
- Best Local LLM on Mac M4 16 GB: Setup, MLX vs GGUF and Real Speed
- Best Local LLM on Mac M4 Pro 48GB: Setup, Quantization and Real Speed
How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks
Xing4.0-29B-A4B is a 31.22 billion parameter mixture-of-experts model from China Telecom's Xing series that only activates 4 billion parameters per token and natively handles a 256K token context. Here is what it costs in VRAM and how to actually serve it.
Xing4.0-29B-A4B is a 31.22 billion parameter open-weight language model from XingChen-AGI (China Telecom's AI division, the team formerly behind TeleChat), released September 16, 2026. It's a mixture-of-experts design that activates only 4 billion parameters per token, natively supports a 256K token context extensible to 512K, and is built for agent work: tool calling, multi-step planning, and coding. The weights are Apache 2.0 licensed.
Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and VRAM requirements against the primary source before recommending an install path.
A new top-20 entrant on Hugging Face's trending list
Xing4.0-29B-A4B entered the top 20 of Hugging Face's overall trending list within two days of release. At time of writing it has 421 likes and 3,073 downloads on the Hugging Face repo, and its predecessor repo, Tele-AI/TeleChat3 on GitHub, has 57 stars. No demo Space is linked yet.
Specs at a glance
| Field | Value |
|---|---|
| Publisher | XingChen-AGI (China Telecom Artificial Intelligence Technology Co.) |
| Parameters | 31.22 billion total (BF16 safetensors, per the HF API); the model card rounds this to 29B total, 4B active per token |
| Architecture | Mixture-of-experts: 64 routed experts + 1 shared expert, 4 active per token, MLA attention, MTP |
| Layers / hidden size | 40 layers, 3,584 hidden size |
| Context length | 256K tokens, extensible to 512K |
| Pipeline | Text generation, tuned for agent and tool-calling workloads |
| License | Apache 2.0 |
| Release date | September 16, 2026 |
| GitHub stars | 57 (Tele-AI/TeleChat3, the project's prior name) |
Hardware: what 31.22B total parameters actually costs you
XingChen-AGI hasn't published a VRAM table, so the honest way to size hardware is to work from the weights themselves. The safetensors checkpoint is almost entirely BF16: 31,215,028,352 of its 31,215,031,088 parameters are BF16, the rest is a handful of F32 buffers. At two bytes per parameter, that's about 62.4GB just to hold the weights. Because this is a mixture-of-experts model, that number doesn't shrink even though only 4 billion parameters activate per token: every expert has to sit in memory in case the router picks it, so the loading cost tracks the 31.22B total, not the 4B active count.
That puts this well outside single consumer-GPU territory. The project's own vLLM launch example sets --tensor-parallel-size 2, which is the clearest signal of what XingChen-AGI expects you to run it on: at least two GPUs with a combined 64GB-plus of VRAM, and meaningfully more once you add a KV cache for anything approaching the 256K context window. This is a workstation or small-cluster model, not a weekend build.
Benchmark scores
XingChen-AGI compares Xing4.0-29B-A4B against Gemma4-26B-A4B and Qwen3.6-35B-A3B, two similarly sized MoE models, on agent and coding benchmarks. The model card reports these scores:
| Benchmark | Xing4.0-29B-A4B | Gemma4-26B-A4B | Qwen3.6-35B-A3B |
|---|---|---|---|
| SWE-bench Verified | 75.00 | 53.00 | 76.00 |
| Terminal-Bench 2.1 | 57.50 | 30.00 | 51.50 |
| SWE-bench Multilingual | 66.00 | 51.00 | 67.20 |
| Claw-Eval | 76.55 | 71.49 | 74.54 |
| DeepresearchBII | 60.80 | 39.30 | 59.70 |
| AIME2026 | 90.00 | 88.30 | 92.70 |
Xing4.0-29B-A4B beats Gemma4-26B-A4B on every row shown, most heavily on Terminal-Bench 2.1 and DeepresearchBII, and runs close to Qwen3.6-35B-A3B, trading the lead depending on the benchmark. Terminal-Bench 2.1, which scores agents operating a real terminal, is where the gap over both competitors is widest.
Serve it with vLLM
The project's GitHub repo documents launch commands for vLLM, SGLang, and KTransformers. The vLLM command, run across two GPUs, is the most standard path once your vLLM build includes the merged support:
Once the server is up, talk to it through the standard OpenAI-compatible client. The model card's own example enables the thinking mode by default and sets the sampling parameters XingChen-AGI recommends:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
completion = client.chat.completions.create(
model="Xing4.0-29B-A4B",
messages=[{"role": "user", "content": "Briefly explain the basic principles of quantum computing."}],
temperature=1.0,
top_p=0.95,
extra_body={
"repetition_penalty": 1.05,
"chat_template_kwargs": {"enable_thinking": True},
},
)
print(completion.choices[0].message.content)FAQ
What is Xing4.0-29B-A4B used for?
Agent and coding workloads: tool calling, multi-step planning, and long-context reasoning, backed by a 256K token context window extensible to 512K. It's built to plug into agent frameworks rather than serve as a general chatbot.
How many parameters does it have?
31.22 billion total parameters, per the safetensors metadata in the Hugging Face API, with 4 billion active per token through its mixture-of-experts routing. The model card rounds this to 29B total, 4B active.
What GPU do I need to run it?
The BF16 weights alone need about 62.4GB, and because it's a mixture-of-experts model, that full amount has to be resident regardless of the 4B active-per-token count. The project's own vLLM example assumes two GPUs (--tensor-parallel-size 2), so plan on 64GB-plus of combined VRAM before adding room for the KV cache.
Can I run it on one consumer GPU?
Not at full precision. At 31.22B total parameters it doesn't fit a single 24GB or even 48GB card in BF16. A quantized build would lower that floor, but XingChen-AGI hasn't published one as of this writing, so budget for a multi-GPU setup or wait for community GGUF conversions.
Is it free to use commercially?
Yes. The weights are licensed under Apache 2.0, with no separate commercial tier or approval step required.
Where to go from here
- Not ready for a multi-GPU build? Run DeepSeek V4 locally: which size actually fits your hardware walks through picking a size that matches what you actually own.
- Want the biggest single-card tier we've covered? How to set up local AI on 24GB of VRAM is the practical ceiling before you need a second GPU.
- Want to compare it against other open models first? Browse the local AI model directory to see specs side by side before you commit disk space to a download.
Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)
Watch related tutorials
9:42
10:30
11:05
9:50
13:30
16:45Weekly local AI drops
New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.