In this categoryLocal AI · 46
More guides
Local AIAdvanced

How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks

Xing4.0-29B-A4B is a 31.22 billion parameter mixture-of-experts model from China Telecom's Xing series that only activates 4 billion parameters per token and natively handles a 256K token context. Here is what it costs in VRAM and how to actually serve it.

7 minAdvanced

Xing4.0-29B-A4B is a 31.22 billion parameter open-weight language model from XingChen-AGI (China Telecom's AI division, the team formerly behind TeleChat), released September 16, 2026. It's a mixture-of-experts design that activates only 4 billion parameters per token, natively supports a 256K token context extensible to 512K, and is built for agent work: tool calling, multi-step planning, and coding. The weights are Apache 2.0 licensed.

Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and VRAM requirements against the primary source before recommending an install path.

Xing4.0-29B-A4B entered the top 20 of Hugging Face's overall trending list within two days of release. At time of writing it has 421 likes and 3,073 downloads on the Hugging Face repo, and its predecessor repo, Tele-AI/TeleChat3 on GitHub, has 57 stars. No demo Space is linked yet.

Specs at a glance

FieldValue
PublisherXingChen-AGI (China Telecom Artificial Intelligence Technology Co.)
Parameters31.22 billion total (BF16 safetensors, per the HF API); the model card rounds this to 29B total, 4B active per token
ArchitectureMixture-of-experts: 64 routed experts + 1 shared expert, 4 active per token, MLA attention, MTP
Layers / hidden size40 layers, 3,584 hidden size
Context length256K tokens, extensible to 512K
PipelineText generation, tuned for agent and tool-calling workloads
LicenseApache 2.0
Release dateSeptember 16, 2026
GitHub stars57 (Tele-AI/TeleChat3, the project's prior name)
License is fully open
Apache 2.0 covers the weights, so there's no separate commercial tier or gated request form to get through before you can use it.

Hardware: what 31.22B total parameters actually costs you

XingChen-AGI hasn't published a VRAM table, so the honest way to size hardware is to work from the weights themselves. The safetensors checkpoint is almost entirely BF16: 31,215,028,352 of its 31,215,031,088 parameters are BF16, the rest is a handful of F32 buffers. At two bytes per parameter, that's about 62.4GB just to hold the weights. Because this is a mixture-of-experts model, that number doesn't shrink even though only 4 billion parameters activate per token: every expert has to sit in memory in case the router picks it, so the loading cost tracks the 31.22B total, not the 4B active count.

That puts this well outside single consumer-GPU territory. The project's own vLLM launch example sets --tensor-parallel-size 2, which is the clearest signal of what XingChen-AGI expects you to run it on: at least two GPUs with a combined 64GB-plus of VRAM, and meaningfully more once you add a KV cache for anything approaching the 256K context window. This is a workstation or small-cluster model, not a weekend build.

Framework support is still landing
The vLLM, SGLang, and KTransformers launch commands published in the project's GitHub repo are attached to pull requests that hadn't merged into those projects' main branches as of this writing. Check the repo for current PR status before you build a serving stack around any one of them.

Benchmark scores

XingChen-AGI compares Xing4.0-29B-A4B against Gemma4-26B-A4B and Qwen3.6-35B-A3B, two similarly sized MoE models, on agent and coding benchmarks. The model card reports these scores:

BenchmarkXing4.0-29B-A4BGemma4-26B-A4BQwen3.6-35B-A3B
SWE-bench Verified75.0053.0076.00
Terminal-Bench 2.157.5030.0051.50
SWE-bench Multilingual66.0051.0067.20
Claw-Eval76.5571.4974.54
DeepresearchBII60.8039.3059.70
AIME202690.0088.3092.70

Xing4.0-29B-A4B beats Gemma4-26B-A4B on every row shown, most heavily on Terminal-Bench 2.1 and DeepresearchBII, and runs close to Qwen3.6-35B-A3B, trading the lead depending on the benchmark. Terminal-Bench 2.1, which scores agents operating a real terminal, is where the gap over both competitors is widest.

Serve it with vLLM

The project's GitHub repo documents launch commands for vLLM, SGLang, and KTransformers. The vLLM command, run across two GPUs, is the most standard path once your vLLM build includes the merged support:

zsh - serve Xing4.0-29B-A4B with vLLM
$vllm serve /path/to/Xing4.0-29B-A4B \
--tensor-parallel-size 2 \
--trust-remote-code \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--reasoning-parser xing4 \
--tool-call-parser xing4
$

Once the server is up, talk to it through the standard OpenAI-compatible client. The model card's own example enables the thinking mode by default and sets the sampling parameters XingChen-AGI recommends:

xing4_client.py
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

completion = client.chat.completions.create(
    model="Xing4.0-29B-A4B",
    messages=[{"role": "user", "content": "Briefly explain the basic principles of quantum computing."}],
    temperature=1.0,
    top_p=0.95,
    extra_body={
        "repetition_penalty": 1.05,
        "chat_template_kwargs": {"enable_thinking": True},
    },
)

print(completion.choices[0].message.content)
Two temperature presets, not one
XingChen-AGI recommends temperature 1.0 / top_p 0.95 for general reasoning and a lower temperature 0.8 / top_p 0.95 for coding and agent tasks, both with repetition_penalty 1.05. Match the preset to the job rather than using one setting for everything.

FAQ

What is Xing4.0-29B-A4B used for?

Agent and coding workloads: tool calling, multi-step planning, and long-context reasoning, backed by a 256K token context window extensible to 512K. It's built to plug into agent frameworks rather than serve as a general chatbot.

How many parameters does it have?

31.22 billion total parameters, per the safetensors metadata in the Hugging Face API, with 4 billion active per token through its mixture-of-experts routing. The model card rounds this to 29B total, 4B active.

What GPU do I need to run it?

The BF16 weights alone need about 62.4GB, and because it's a mixture-of-experts model, that full amount has to be resident regardless of the 4B active-per-token count. The project's own vLLM example assumes two GPUs (--tensor-parallel-size 2), so plan on 64GB-plus of combined VRAM before adding room for the KV cache.

Can I run it on one consumer GPU?

Not at full precision. At 31.22B total parameters it doesn't fit a single 24GB or even 48GB card in BF16. A quantized build would lower that floor, but XingChen-AGI hasn't published one as of this writing, so budget for a multi-GPU setup or wait for community GGUF conversions.

Is it free to use commercially?

Yes. The weights are licensed under Apache 2.0, with no separate commercial tier or approval step required.

Where to go from here

Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)

Watch related tutorials

Free weekly email

Weekly local AI drops

New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.

Tags
#xing4.0-29b-a4b#xingchen-agi xing4#run xing4 locally#telechat xing series#mixture of experts local llm#256k context model