In this categoryLocal AI · 46
- How to Run Audio8 ASR Infinite Locally: Streaming Speech Recognition That Never Stops
- How to Run TeleOCR Locally: One Model for Scanned and Photographed Documents
- How to Run Xing4.0-29B-A4B Locally: A 256K-Context MoE Model for Agent Tasks
- How to Run IndicF5 Locally: Voice Cloning for 11 Indian Languages
- How to Run Fish Audio S2 Pro Locally: Multilingual TTS with Inline Emotion Control
- How to Run ZDTaichu5.0-9B Locally: A 9.79B Vision-Language Model Built for Spatial Reasoning
- How to Run Qwen3-TTS-12Hz-1.7B-VoiceDesign Locally: Voice Design From a Text Description
- How to Run YuE2-3B Locally: Open Source Music Generation That Beats Suno v5 on Benchmarks
- MiniCPM5-2B: Install and Run OpenBMB's 2B Model That Outscores 4B Rivals
- BERT Base Uncased: Specs, How to Run It, and Why It's Still Trending in 2026
- How to Run tencent/AuK Locally: Zero-Shot TTS, Voice Cloning and Speech Editing
- Stable Diffusion 3.5 Medium: What's Different From SDXL and Flux
- Qwen-Image-2512: The Text-to-Image Model That Renders Real Text
- How to Run Irodori-TTS Anime Locally: Japanese Voice Cloning and Voice Design
- How to Run Breeze TTS 2 Locally: Voice Clone, Voice Design and Voice Direction
- How to Run Kokoro-82M Locally: the Fastest Open-Weight TTS Model
- How to Install Ollama on macOSStart
- How to Install Ollama on Windows
- How to Run Llama 3 Locally with Ollama
- How to Pick the Right Local AI Model for Your Hardware
- Best GGUF Models to Run by VRAM Tier (8GB, 12GB, 16GB, 24GB, 48GB)
- Run LLMs in Your Browser With WebGPU: No Install, No Server (WebLLM)
- How to Use Ollama as a Drop-In OpenAI API
- GGUF vs MLX vs NVFP4: Local AI Quantization Formats Explained
- Best GPU for Running AI Locally in 2026Start
- How Much RAM Do You Need for Local AI?
- Mac vs PC for Local AI: Which Should You Choose?
- How to Build a Local AI Workstation on Any Budget
- How to Set Up Local AI on 8GB of VRAM
- How to Set Up Local AI on 12GB of VRAM
- Best Local LLM for 16 GB VRAM: Setup, Quantization and Real Speed
- How to Set Up Local AI on 24GB of VRAM
- Best Local LLM on Mac M4 16 GB: Setup, MLX vs GGUF and Real Speed
- Best Local LLM on Mac M4 Pro 48GB: Setup, Quantization and Real Speed
How to Run TeleOCR Locally: One Model for Scanned and Photographed Documents
TeleOCR is a 1.42 billion parameter open-weight vision-language model from StarDoc-AI that parses both digital PDFs and camera-photographed documents with the same weights. Here is what it scores on, what it costs in VRAM, and how to run it.
TeleOCR is a 1.42 billion parameter open-weight vision-language model from StarDoc-AI that reads both digital PDFs and camera-photographed documents with one set of weights. Released August 14, 2026 under Apache 2.0, it scores 96.87 overall on OmniDocBench v1.6, ahead of similarly sized OCR models, and installs with a single pip command.
Written by Priya Raghunathan, local-AI hardware reviewer. I check parameter counts and hardware requirements against the primary source before recommending an install path.
A new top-20 entrant on Hugging Face's trending list
TeleOCR entered the top 20 of Hugging Face's overall trending list. At time of writing it has 590 likes and 27,837 downloads on the Hugging Face repo, plus 238 stars on the caipeng328/TeleOCR GitHub repo it ships alongside, and a demo Space, StarDoc-AI/navidc-ocr-demo, with 51 likes if you want to try a document before installing anything.
Specs at a glance
| Field | Value |
|---|---|
| Publisher | StarDoc-AI |
| Parameters | 1.42 billion (1,415,072,768 in BF16, per the HF API's safetensors metadata) |
| Pipeline | Image-text-to-text: a document image plus a task prompt in, structured text out |
| Architecture | Built on Qwen2.5-VL (Qwen2_5_VLForConditionalGeneration) |
| Languages | Chinese and English |
| Document types | Digital documents and camera-captured (photographed, distorted) documents in one framework |
| License | Apache 2.0 |
| Release date | August 14, 2026 |
| GitHub stars | 238 (caipeng328/TeleOCR) |
What makes it different: one model for two kinds of documents
Most open OCR models pick a lane. Some are tuned for clean digital PDFs, others for photos of paper taken at an angle with a phone camera. TeleOCR targets both from a single checkpoint. Its training pipeline includes geometry-aware modeling for warped or tilted pages, a sampling method called Curvature-Guided Douglas-Peucker Sampling for tracing distorted text lines, and a self-verification step where the model checks its own output by rendering it back into an image and comparing. StarDoc-AI also built the training labels with a consensus-voting setup across multiple nodes rather than relying on a single labeler, and trained tables and formulas with a separate content-and-structure objective so a table's layout and its cell text are learned as related but distinct problems.
Benchmarks: how it stacks up against similarly sized models
On OmniDocBench v1.6, the standard benchmark for document parsing, TeleOCR posts the highest overall score among the specialized VLMs it is compared against, and does it with the same or fewer parameters.
| Model | Parameters | Overall | Table TEDS | Read Order Edit (lower is better) |
|---|---|---|---|---|
| TeleOCR | 1.2B | 96.87 | 97.05 | 0.122 |
| OvisOCR2 | 0.8B | 96.58 | 94.76 | 0.111 |
| PaddleOCR-VL-1.6 | 0.9B | 96.33 | 94.76 | 0.127 |
| MinerU2.5-Pro | 1.2B | 95.75 | 93.42 | not published in this table |
TeleOCR also placed first on the ICDAR 2026 Sci-ImageMiner benchmark, a separate contest for extracting structured data out of scientific figures, and StarDoc-AI reports it beating MinerU 2.5 Pro and PaddleOCR-VL 1.6 on the EMNLP 2026 Dr.DocBench Challenge (67.96 overall versus 62.26 and 55.11).
Hardware: what 1.42B parameters costs you
StarDoc-AI hasn't published a VRAM table, so the honest way to size hardware is to work from the weights themselves. The safetensors checkpoint stores 1.42 billion parameters in BF16, which is 2 bytes per parameter, putting the weights at roughly 2.8GB. On top of that, a vision-language model spends extra VRAM turning each input image into vision tokens before it ever generates text, so budget headroom beyond the bare weight size. A 6GB card should run it, and an 8GB card runs it comfortably with room for a KV cache on longer documents.
If you're deciding what to build a local-AI box around, see our guide on setting up local AI on 8GB of VRAM, a tier TeleOCR runs on with headroom for other models alongside it.
Install it
TeleOCR loads through the transformers library. There is no separate SDK to install.
Run your first parse
Load the processor and model with trust_remote_code enabled, then hand an image and a task prompt to an infer() function. The same function handles plain text extraction, table extraction in OTSL format, LaTeX formulas, code snippets, and layout analysis, just by changing the prompt string.
from transformers import AutoProcessor, AutoModel
from PIL import Image
import torch
processor = AutoProcessor.from_pretrained(
"StarDoc-AI/TeleOCR", trust_remote_code=True, use_fast=True
)
model = AutoModel.from_pretrained(
"StarDoc-AI/TeleOCR", trust_remote_code=True, torch_dtype=torch.bfloat16,
).cuda().eval()
def infer(image: Image.Image, prompt: str) -> str:
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [
{"type": "image"},
{"type": "text", "text": prompt},
]},
]
chat_prompt = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True,
)
inputs = processor(
text=[chat_prompt], images=[image.convert("RGB")],
padding=True, return_tensors="pt",
).to(device=model.device, dtype=model.dtype)
output_ids = model.generate(
**inputs, use_cache=True, max_new_tokens=4096, do_sample=False,
)
output_ids = output_ids.cpu().tolist()[0][len(inputs.input_ids[0]):]
return processor.batch_decode(
[output_ids], skip_special_tokens=True, clean_up_tokenization_spaces=False,
)[0].strip()
image = Image.open("invoice.png")
raw_text = infer(image, "Please output the text content from the image.")
raw_otsl = infer(image, "This is the image of a table. Please output the table in OTSL format.")
raw_formula = infer(image, "Please write out the expression of the formula in the image using LaTeX format.")
print(raw_text)Batch processing: the vLLM path in the GitHub repo
The single-image script above is fine for trying TeleOCR out, but it's not what you want for a folder of hundreds of scans. The caipeng328/TeleOCR GitHub repo ships a separate batch path built on vLLM, run through an infer.py script rather than a single-image call. Clone the repo and install it in editable mode first.
python infer.py \
--image_sub_path "${IMAGE_SUB_PATH}" \
--result_save_path "${RESULT_SAVE_PATH}" \
--use_async \
--override \
model_path="StarDoc-AI/TeleOCR" \
BACKEND="vllm-async-engine" \
LAYOUT_MODE="Detection"BACKEND switches between a synchronous vllm-engine and the async vllm-async-engine shown above. MAX_MODEL_LEN and GPU_MEMORY_UTILIZATION control how much of the GPU vLLM is allowed to claim, and PDF_TOOLS lets you pick PyMuPDF or pypdfium2 for converting PDFs to page images before they hit the model. None of this is an OpenAI-compatible server, it's a batch script, so plan on point-in-time runs over a folder rather than a long-lived endpoint.
A community member also converted the earlier NaviDC-OCR weights to GGUF for llama.cpp, published at nandraj/NaviDC-OCR-GGUF on Hugging Face. It predates the TeleOCR rename, so check that repo directly if you need a CPU-friendly quantized build rather than assuming it tracks the latest release.
FAQ
What is TeleOCR used for?
Parsing documents into structured text: plain text extraction, table extraction in OTSL format, LaTeX formulas, code snippets, and layout analysis, from either clean digital pages or photos of paper documents.
How many parameters does it have?
1.42 billion parameters, per the safetensors metadata in the Hugging Face API (1,415,072,768 stored in BF16). The model card itself describes it as roughly 1.2 billion parameters, a lightweight VLM built on Qwen2.5-VL.
What GPU do I need to run it?
The BF16 weights are about 2.8GB, so a 6GB card can run it and an 8GB card runs it with headroom. StarDoc-AI hasn't published an official VRAM figure, so budget some extra for vision token processing on top of the base weight size.
Is TeleOCR the same as NaviDC-OCR?
Yes. StarDoc-AI renamed NaviDC-OCR to TeleOCR on September 10, 2026. Future releases continue under the TeleOCR name, but older links and the community GGUF conversion still reference NaviDC-OCR.
Can I use it commercially?
Yes. Both the code and the model weights are licensed under Apache 2.0, with no separate paid tier required for commercial use.
Where to go from here
- Building a local-AI box around a small model like this? How to set up local AI on 8GB of VRAM covers a hardware tier TeleOCR runs on with room to spare.
- Want another self-hosted OCR option to compare against? Self-host baidu Unlimited-OCR covers a different open-weight document parser, whole-PDF-in-one-pass rather than image-by-image.
- Want to compare it against other open OCR and vision-language models first? Browse the local AI model directory to see specs side by side before you commit disk space to a download.
Local models run better with more VRAM. CompareRTX GPUs on Amazonbefore you upgrade.(affiliate link. We may earn a commission at no extra cost. Disclosure)
Watch related tutorials
9:42
10:30
11:05
9:50
13:30
16:45Weekly local AI drops
New models, what runs on your hardware, and the guides to set them up. One email a week, unsubscribe any time.