A 1.2B open-source vision-language model unifies digital and camera-captured document parsing, topping OmniDocBench and beating larger competitors like MinerU 2.5 Pro.

Full article content could not be extracted automatically. Read the original below.