PaddleOCR-VL-1.6, a compact 0.9B parameter vision-language model, has set a new state-of-the-art on the OmniDocBench v1.6 benchmark with 96.3% accuracy. The model demonstrates significant improvements in text, formula, and table recognition, while also enhancing capabilities for challenging scenarios like ancient texts, rare characters, seals, and complex charts. A key feature is its ability to output structured results in both Markdown and JSON formats, which simplifies integration into document processing workflows. For developers working on OCR, document digitization, or automated data extraction, this release offers a strong open-source alternative to commercial solutions. The model's compact size makes it feasible for deployment in resource-constrained environments, potentially lowering the barrier for high-quality document understanding in production systems.
PaddleOCR-VL-1.6 achieves 96.3% on OmniDocBench, improving text, formula, and table recognition with structured output support.