Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions
A survey paper on arXiv reviews recent developments integrating commonsense knowledge into computer vision tasks, covering approaches based on knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. It outlines limitations related to dataset bias and knowledge incompleteness.
The paper presents a comprehensive survey of recent developments that integrate commonsense knowledge into computer vision tasks. It systematically reviews approaches based on knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. It also outlines current limitations related to dataset bias and knowledge incompleteness.
The survey highlights a shift from CNN-based object identification to holistic scene interpretation using commonsense knowledge, enabling models to reason about relationships among objects and actions. Key technical approaches include knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers.
Integrating commonsense reasoning into computer vision could improve AI systems' ability to interact meaningfully with humans and the environment, potentially leading to more precise predictions in real-world applications such as robotics, autonomous systems, and human-computer interaction.
Enhanced commonsense reasoning in computer vision can lead to more robust AI products in areas like autonomous driving, smart surveillance, and assistive technologies, potentially reducing errors and improving user trust.
Future work may address dataset bias and knowledge incompleteness, and further develop neuro-symbolic and transformer-based methods to enhance commonsense reasoning in vision. Observable next signals include new benchmarks and models that demonstrate improved spatial and contextual reasoning.