Tencent open-sourced WeVisDoc-4B document parser
Tencent released WeVisDoc-4B, an end-to-end document parser fine-tuned from Qwen3-VL-4B-Instruct, on Hugging Face under Apache-2.0. It achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across PureDocBench tracks.
China context
- Original name
- 腾讯
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Builders can self-host WeVisDoc-4B under Apache-2.0 to convert page images into structured Markdown, avoiding per-page API costs and data leaving their infrastructure.
- For investors
- Tencent's release of a top-ranked open-weights document parser intensifies competition in the document AI market, potentially compressing margins for commercial OCR and document processing startups.
WeVisDoc-4B is an end-to-end document parser for page images, fine-tuned from Qwen3-VL-4B-Instruct. It converts pages into structured Markdown with LaTeX formulas and HTML tables. On OmniDocBench v1.6, it scores 95.38 Overall, and on PureDocBench it averages 75.54 across Clean, Digital, and Real tracks, ranking first among compared end-to-end parsers in all four settings. The model is available on Hugging Face under Apache-2.0 license.
WeVisDoc-4B is a 4B-parameter vision-language model fine-tuned from Qwen3-VL-4B-Instruct for document parsing. It outputs structured Markdown with LaTeX formulas and HTML tables. Evaluation results are means over three inference runs. The model card provides quick-start instructions using vLLM 0.11.1 or local Transformers inference, with configurable context length (default 32768 tokens) and GPU memory utilization.
Developers outside China can now use a state-of-the-art open-weights document parser under Apache-2.0, reducing the cost and complexity of building document conversion pipelines compared to proprietary OCR APIs. This directly pressures commercial document AI vendors by offering a free, self-hostable alternative with top benchmark scores.
For businesses, WeVisDoc-4B offers a cost-effective, open-source solution for converting scanned documents and PDFs into structured Markdown, potentially reducing reliance on paid OCR services. Its Apache-2.0 license permits commercial use, enabling integration into document management, data extraction, and RAG pipelines.
Next signals to check: whether Tencent releases a technical report or paper detailing training data and methodology; whether the model is integrated into cloud services or enterprise products; and whether independent evaluations confirm the reported benchmark scores.