MindTopo: Can Foundation Models Reason in Topological Space?
MindTopo is a benchmark of topological intuition across five properties: continuity, separation, order, enclosure, and knots. It evaluates each property at two cognitive levels: reasoning and planning. The benchmark contains 11,030 instances across 13 procedurally generated task types with controllable difficulty. 14 MLLMs were benchmarked, and agent configurations augmented with image and video generation were studied, including 3 video generative models in planning settings. Every MLLM performed better on reasoning than on planning, and the best-performing model remained far below observed human performance.
Researchers introduced MindTopo, a benchmark designed to assess topological reasoning in foundation models. The benchmark covers five topological properties—continuity, separation, order, enclosure, and knots—and evaluates models at reasoning and planning levels. It includes 11,030 instances across 13 procedurally generated task types. The study benchmarked 14 multimodal large language models (MLLMs) and explored agent configurations with image and video generation, including three video generative models in planning settings. Results showed that all MLLMs performed better on reasoning tasks than on planning tasks, and even the best model fell far short of human performance.
The benchmark separates reasoning (identifying topological relations or inferring changes) from planning (closed-loop agent policy selecting environment actions). The inclusion of video generative models in planning settings suggests an exploration of world-model-like capabilities. The consistent reasoning-over-planning performance gap indicates that current MLLMs struggle to translate topological understanding into action sequences, a key challenge for embodied AI.
This work highlights a specific weakness in spatial reasoning that could affect applications in robotics, navigation, and simulation. The use of procedurally generated tasks with controllable difficulty provides a scalable evaluation method. The gap between reasoning and planning performance may signal that current models are not yet ready for complex real-world spatial tasks requiring topological understanding.
For companies developing spatial AI or robotics, MindTopo offers a way to assess and compare model capabilities in topological reasoning. The identified performance gap suggests a market opportunity for models that can bridge reasoning and planning. The benchmark's procedural generation allows for cost-effective, scalable evaluation.
Future work may focus on improving planning capabilities through better integration of generative models or reinforcement learning. The benchmark could become a standard for evaluating topological reasoning in foundation models. Progress on MindTopo may correlate with advances in embodied AI and robotics.