Tencent open-sourced UI-Mate-democua-27B, a demonstration-guided GUI agent
Tencent released UI-Mate-democua-27B, a 27B demonstration-guided GUI agent checkpoint, on Hugging Face under Apache-2.0. It is built on Qwen3.6-27B and accepts task instructions, screenshots, interaction history, and an optional demonstration workflow.
China context
- Original name
- 腾讯
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Builders outside China can download the Apache-2.0 weights from Hugging Face and integrate the model into GUI automation pipelines using the provided OpenAI-compatible interface and pyautogui-compatible actions.
- For investors
- Investors can monitor adoption of UI-Mate checkpoints and the UI-Mate repository as an indicator of Tencent's competitiveness in open-weight GUI agents, a segment where open models may pressure proprietary computer-use APIs.
UI-Mate-democua-27B is a demonstration-guided checkpoint of UI-Mate, an open-weight foundation GUI agent. It observes live screenshots, reasons over the visible state, and produces structured keyboard and mouse actions. It can take one recorded workflow as input and carry that procedure over to a new task. The model is trained with supervised fine-tuning on a mixture of general computer-use data and demonstration-augmented data, preserving instruction-only competence while adding demonstration-guided execution. A demonstration is treated as guidance rather than a fixed action script; recorded coordinates are never replayed, and the model re-plans from the live interface when content, layout, or application state differs. The model has 27B parameters, uses Qwen3.6-27B as its base, and is licensed under Apache-2.0. It is an agent checkpoint rather than a standalone visual-chat model, and Tencent recommends using the official prompt, response parser, and interaction harness from the UI-Mate repository.
The model uses a demonstration-guided execution approach where a recorded workflow is normalized, annotated by a vision-language model along four axes (screen state, intent, action taken, target location), segmented into named subtasks with checkable completion criteria, and supplied at inference as a compact view of the active subtask. Training deliberately withholds part of the guidance (key actions like focus clicks and scrolling are omitted) so the model must infer missing steps from screenshots. The training mixture covers full alignment, partial misalignment, and irrelevance between guidance and screen, with full alignment as the majority case. At inference, the complete action sequence of the current subtask is passed without key-action extraction. Evaluation on OSWorkerBench-Subset (33 tasks) shows strict success improving from 17.17 to 35.35 (+18.18 pp) and progress from 67.85 to 81.14 (+13.29 pp) when adding one demonstration. On OSWorld-Subset (30 tasks), progress improves from 40.27 to 65.75 (+25.48 pp). On GameDev (10 tasks), average score improves from 76.76 to 81.15 (+4.39 pp).
Developers outside China can now use an Apache-2.0 licensed 27B GUI agent that improves task success by up to 18.18 percentage points when given a single demonstration, reducing the need for extensive prompt engineering or fine-tuning for computer-use tasks.
The model enables one-shot procedural learning from a single demonstration, which can lower the cost and time to automate GUI workflows. Its Apache-2.0 license allows commercial use, and its structured actions are compatible with pyautogui and served through an OpenAI-compatible interface, easing integration.
Observable next signals include whether Tencent releases the UI-Mate-27B and UI-Mate-9B checkpoints mentioned in the model card, whether the UI-Mate repository gains adoption among developers building GUI automation tools, and whether independent benchmarks confirm the reported improvements.