GPT-5.5, Claude Fable 5 · Jul 17, 2026

An Exam for Active Observers

The paper 'ActiveVision' releases a benchmark containing 17 tasks designed to measure the active observation ability of multimodal large language models. Evaluation results show that GPT-5.5, at the highest reasoning effort level, solves only 10.6% of tasks and scores zero on 11 of the 17 tasks; Claude Fable 5 solves only 3.5%; while three human participants average 96.1%. Even when models write and run their own visual code, the gap persists.

What happened

The new benchmark ActiveVision reveals that cutting-edge multimodal large language models are severely deficient in active observation ability, with GPT-5.5 and Claude Fable 5 performing far worse than humans.

Technical significance

Current multimodal large language models lack a human-like closed-loop mechanism for active observation, unable to dynamically adjust visual attention based on intermediate hypotheses, leading to failure on tasks requiring repeated visual perception. Even with code generation capabilities, models still struggle to reliably process real images and cannot autonomously discover code errors, indicating fundamental limitations in visual reasoning.

Industry impact

This benchmark exposes the vulnerability of multimodal models in real-world visual tasks, potentially affecting deployment confidence in applications that rely on active vision (e.g., robotics, autonomous driving, medical imaging). Model providers need to rethink visual architectures, shifting from static description to dynamic perception.

What to watch

Observable signals include: whether model providers release architectures or training methods that improve active vision; whether subsequent research adopts ActiveVision as a standard evaluation; and whether active vision capability becomes a new dimension of competition for multimodal models.

Decision value

This research identifies a key shortcoming of current multimodal AI products, potentially driving investment in active vision technology and providing selection references for enterprise applications requiring reliable visual understanding (e.g., quality inspection, surveillance).

Evidence