{"name":"chinese-ocr","title":"Chinese OCR","tables":[{"heading":"Precision: OmniDocBench v1.6 overall, out of 100","columns":["Model","Lab","Size","Score","Source"],"rows":[["TeleOCR","China Telecom AI","1.2B","96.91","Official leaderboard"],["OvisOCR2","Alibaba (ATH-MaaS)","0.8B","96.47","Official leaderboard"],["PaddleOCR-VL-1.6","Baidu PaddlePaddle","0.9B","96.34","Official leaderboard"],["MinerU2.5-Pro","OpenDataLab / Shanghai AI Lab","1.2B","95.75","Official leaderboard"],["GLM-OCR","Zhipu (Z.ai)","0.9B","95.22","Official leaderboard"],["HunyuanOCR-1.5","Tencent","1B","94.74","Self-reported or another lab's table"],["Unlimited-OCR","Baidu","3B MoE (0.5B active)","93.92","Self-reported or another lab's table"],["Qianfan-OCR","Baidu Qianfan","4B","93.90","Self-reported or another lab's table"],["dots.ocr","Xiaohongshu (rednote)","~3B","90.77","Self-reported or another lab's table"],["DeepSeek-OCR 2","DeepSeek","3B MoE (~0.5B active)","90.25","Official leaderboard"],["Qwen3-VL-235B","Reference: large general VLM","235B","89.78","Reference, not a contender"],["Not on v1.6","Qwen3-VL-4B (about 86.8 on v1.5), the qwen-vl-ocr API, classic PaddleOCR","","n/a",""]],"note":"Full-page parsing of text, tables, formulas and reading order. Gaps under about one point are noise."},{"heading":"Speed: pages per second, self-hosted","columns":["Model","Size","Pages per second","How it was measured"],"rows":[["OvisOCR2","0.8B","9.80","~9.8 p/s peak · RTX 5090 · 32 concurrent · third-party test"],["DeepSeek-OCR 2","3B MoE (~0.5B active)","4.65","~4.65 p/s · A100 · third-party test of v1 (~200k pages/day)"],["MinerU2.5-Pro","1.2B","2.12","2.12 p/s · A100 · self-reported"],["PaddleOCR classic","Small, non-VLM","2.00","~2 p/s (~120 pages/min) · RTX 3090 · third-party"],["GLM-OCR","0.9B","1.86","1.86 PDF p/s · 1 request at a time · Zhipu"],["PaddleOCR-VL-1.6","0.9B","1.43","1.43 p/s · A100 · Baidu (measured on v1.5, same architecture)"],["Qianfan-OCR","4B","1.02","1.02 p/s at 8-bit (0.50 at 16-bit) · A100 · Baidu"],["HunyuanOCR-1.5","1B","0.71","0.71 p/s with DFlash (0.33 without) · H20 · 1 request · Tencent"]],"note":"Each lab measured on its own hardware and concurrency, so treat this as rough order only."},{"heading":"Price: hosted API, USD per million tokens (input / output)","columns":["Model","Price","Where"],"rows":[["GLM-OCR","$0.03 / $0.03","Z.ai"],["DeepSeek-OCR 2","$0.03 / $0.03","Novita AI (split unconfirmed); DeepInfra serves v1 at $0.05 blended"],["qwen3.5-ocr / qwen-vl-ocr","$0.043 / $0.072","qwen-vl-ocr, US/Global region ($0.07/$0.16 Singapore; qwen3.5-ocr $0.069/$0.275)"],["Qwen3-VL-4B","$0.1 / $0.6","llm-stats listing"],["PaddleOCR-VL-1.6","$1 per 1,000 pages","OpenParser (third-party); Baidu Cloud API price not published"],["OvisOCR2","Flat plan","Featherless AI, from $25 a month"],["MinerU2.5-Pro","Not published","mineru.net API; 1,000 priority pages a day per account"],["Qianfan-OCR, Unlimited-OCR","Not published","Baidu APIs (China)"],["TeleOCR, HunyuanOCR, dots.ocr","No paid API","Web app or demo only"]],"note":"Every model in the precision table can also be downloaded and run on your own GPUs."},{"heading":"Bounding boxes: the finest level each returns","columns":["Model","Finest boxes","Details"],"rows":[["PaddleOCR classic","Word, character and table cell","Word boxes, character coordinates, line polygons, table-cell boxes"],["PaddleOCR-VL-1.6","Text line","Layout blocks (pixels, rect/quad/poly) + text-line quads in \"Spotting:\" mode (0–1000)"],["HunyuanOCR-1.5","Text line","Text-line boxes as JSON, 0–1000, in reading order"],["qwen3.5-ocr / qwen-vl-ocr","Text line","Text lines: 4-corner pixel quads + rotated rect (\"advanced_recognition\")"],["TeleOCR","Layout block","Layout blocks, 0–999 grid; optional polygon mode"],["MinerU2.5-Pro","Layout block","Blocks in VLM mode; real line/span boxes only in the classic \"basic\" backend"],["GLM-OCR","Layout block","Layout blocks only (0–1000 in SDK); no line, word or cell boxes"],["Unlimited-OCR","Layout block","Every block boxed, 0–1000, with page separators"],["Qianfan-OCR","Layout block","Blocks via Layout-as-Thought or layout prompt (0–999); 25 categories"],["dots.ocr","Layout block","Blocks in absolute pixels with 11 categories, in JSON"],["DeepSeek-OCR 2","Layout block","Blocks with \"<|grounding|>\" prompt (0–999)"],["Qwen3-VL-4B","Prompted grounding","General grounding as JSON (0–1000), whatever level you ask for"],["OvisOCR2","Figures only","Only charts/images get boxes (0–1000); no text boxes"]]},{"heading":"Input and pages per call","columns":["Model","What it accepts","PDF","Pages the model sees per call","Hosted limits"],"rows":[["qwen3.5-ocr / qwen-vl-ocr","Images; qwen3.5-ocr also PDF (Responses API)","yes (qwen3.5-ocr)","PDF ≤50 pages (parsing), ≤10 other tasks","PDF ≤100 MB; images ≤20 MB"],["Unlimited-OCR","Images; README renders PDFs","via script","Many: tested to 40+ pages","Baidu AI Cloud OCR (price not published)"],["Qwen3-VL-4B","Images, multi-image, video","no","Multi-image","Alibaba Model Studio and others"],["Qianfan-OCR","Images; several pages per request possible","no","Several possible; official skill sends 1","Baidu Qianfan API (China site); OpenRouter lists Qianfan-OCR-Fast but no provider serves it"],["TeleOCR","Images; repo script renders PDFs","via script","1","Web app only (TeleAI)"],["OvisOCR2","Images only; no PDF tool shipped","no","1","Featherless (flat plans from $25/mo)"],["PaddleOCR-VL-1.6","Images; SDK renders PDFs","via SDK","1 (crops batched across pages)","Baidu Cloud API: PDF ≤500 pages, ≤100 MB. Also OpenParser, and Fireworks (dedicated GPU only)"],["MinerU2.5-Pro","PDF, images, Office files","yes (tool)","1","mineru.net: ≤200 pages, ≤200 MB per file; 1,000 priority pages/day"],["GLM-OCR","PDF, JPG, PNG (SDK and API)","yes (SDK/API)","1","Z.ai API: ≤50 MB PDF, ≤100 pages (one doc says 30)"],["HunyuanOCR-1.5","Images (one per call in official client)","no","1","Demo only"],["dots.ocr","Images and PDF (parser script)","via script","1","Demo only"],["DeepSeek-OCR 2","Images; script renders PDFs","via script","1","Novita AI"],["PaddleOCR classic","PDF and images; runs on CPU","yes (tool)","1","Self-host"]]}],"log":[{"date":"261004","line":"New board. 12 Chinese OCR models and Baidu's classic pipeline, compared for precision, speed, price, boxes and input."}],"updatedAt":"2026-10-04T01:04:15.086Z"}