GLM-4.6V Open Source: Zhipu Brings Visual Understanding Directly into Tool Calling
Zhipu released and open-sourced GLM-4.6V and the 9B GLM-4.6V-Flash in December 2025.
Multimodal models advance from image recognition and Q&A to vision-driven tool use, as Zhipu attempts to connect screenshots, documents, charts, and interface understanding into executable agent workflows.
GLM-4.6V adds native multimodal Function Calling within a 128K context, supporting images, screenshots, and document pages as tool inputs, and interpreting visual results returned by tools.
The competitive focus of visual agents shifts from recognition accuracy to interface understanding, action planning, and tool execution outcomes; the lightweight version also expands local deployment and low-latency scenarios.
Suitable for piloting in rollback-friendly document and backend processes; explicit human confirmation is still required for operations involving payments, deletions, or external releases.
Observe GUI operation success rates, visual prompt injection defenses, long-chain document tasks, and the real throughput of the 9B version on edge devices.