Event date · · Ant Group

Ant Group releases PolyOCR-Venus OCR foundation models

FACT STATEMENT

Ant Group's GuangJian Team released PolyOCR-Venus, a family of 2B and 9B OCR foundation models built on Qwen3.5, with a technical report and OCRBench evaluation code. The release includes benchmark results on OCRBench v2.1, CC-OCR, and OmniDocBench v1.6, but model weights and training code are not included.

China context

Original name
蚂蚁集团
Outside China
Not stated in the sources yet
Claims
Company-reported; not yet independently evaluated
For builders
Builders outside China cannot integrate PolyOCR-Venus because model weights are not released; only evaluation code is available.
For investors
Investors should note that Ant Group is advancing OCR foundation models but has not yet open-sourced weights, limiting immediate commercial impact.
What happened

PolyOCR-Venus is a family of 2B and 9B OCR foundation models built on the Qwen3.5 vision-language backbone. It unifies text recognition, localization, document parsing, information extraction, multilingual transformation, and OCR-centric reasoning through a shared instruction-following interface. The models were trained on approximately 60 million SFT instances using a three-phase curriculum and Competence-Guided Policy Optimization (CGPO), which combines verifier-based GRPO with adaptive on-policy distillation. The release includes a technical report and OCRBench evaluation code, but model weights, training code, and CC-OCR/OmniDocBench pipelines are not included.

Technical significance

PolyOCR-Venus uses a Qwen3.5 vision-language backbone and a training framework with curriculum SFT and CGPO. The 9B model achieves 80.42 on OCRBench v2.1 EN, 78.87 on OCRBench v2.1 ZH, 82.43 on CC-OCR, and 91.57 on OmniDocBench v1.6. A layout-guided pipeline with PP-DocLayoutV3 improves OmniDocBench v1.6 to 95.48. The OCRBench v2.1 protocol is a revised version with audited annotations and task-aligned scoring, so scores are not directly comparable to original OCRBench v2.

Industry impact

Developers outside China cannot yet use PolyOCR-Venus because model weights are not released; they can only evaluate the OCRBench code and technical report. This limits immediate adoption and leaves the competitive position of existing OCR models unchanged until weights are published.

Decision value

PolyOCR-Venus could reduce OCR development costs for enterprises if weights are released, but currently the lack of weights means no direct business value outside Ant Group. The technical report and evaluation code may help researchers benchmark their own OCR systems against the reported scores.

What to watch

The next verifiable signal is whether Ant Group releases model weights or API access for PolyOCR-Venus. The repository currently provides only evaluation code, so monitoring the GitHub repository for weight files or an API announcement would confirm availability.

Latest in Chinese AI

  1. TencentReportedly: Tencent Cloud open-sources TeamAI, a Git-based tool for sharing skills across agents
  2. AlibabaReportedly: Alibaba's Qoder Adds Fast Mode With 3-Second First Responses and Lower Credit Use
  3. ShengShu AIReportedly: ShengShu AI releases Vidu Q4 Preview video generation model with 2K and 4K output
  4. AlibabaReportedly: Underdog AI releases Saluki 27B: a 2-bit quantized Qwen3.8-27B at 7.89 GB
  5. ByteDanceReportedly: ByteDance's Doubao pricing strategy influences US AI price competition

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.