TeleOCR targets PDFs and distorted photos with one 1.2B model
StarDoc-AI has released the TeleOCR model card, a roughly 1.2-billion-parameter vision-language model for document parsing. Built on Qwen2.5-VL and released under Apache-2.0, the model handles born-digital pages, scans, and camera captures within one checkpoint. That design can reduce the number of models required to process warped pages, tables, formulas, and complex layouts.
Camera-document pipelines commonly rectify perspective, detect layout regions, run optical character recognition, and invoke separate table or formula parsers. TeleOCR generates content and page geometry directly from the source image, including multi-point layout polygons that can follow curved or skewed regions. Tables are serialized in Open Table Structure Language, or OTSL, while formulas are returned as LaTeX.
1.2B parameters, stronger reported scores
The project reports leading results on three document-parsing suites, including benchmarks designed around degraded and photographed pages. Scores below use the project’s 100-point presentation, with higher values indicating better performance.
The OmniDocBench leaderboard also places TeleOCR ahead of specialized models including OvisOCR2, PaddleOCR-VL-1.6, and MinerU2.5-Pro. TEDS stands for Tree Edit Distance-based Similarity, a metric that compares the structure and content of predicted tables with reference tables. TEDS-S concentrates on structural accuracy.