arXiv:2609.05257cs.AI2026-09

让视觉模型理解日常常识,提升真实场景下的推理能力。

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

  • 融合知识图谱与场景图,实现视觉与常识的跨模态连接。
  • 相比传统模型,对物体关系和上下文理解更准确。
  • 适合研究智能视觉系统、人机交互的学者参考。

计算机视觉中的常识推理旨在融合视觉数据与上下文知识,对提升AI对日常场景的理解至关重要。这不仅增强机器学习模型的表现,还使其能更自然地与人类和环境互动。与仅关注图像内物体识别的CNN模型不同,融入常识知识使模型能够以整体方式解读场景,从而提升对物体间关系与动作的空间推理能力。该整合不仅改善物体识别,还促进对上下文因素的深层理解,最终实现更精确的现实应用预测与交互。本文系统综述了将常识知识融入计算机视觉任务的最新进展,涵盖基于知识图谱、场景图、神经符号模型及常识增强型Transformer的方法。同时指出当前在数据集偏差、知识不完整性和集成挑战方面的局限。最后展望了跨模态推理、可扩展常识注入及神经符号混合架构等未来研究方向,以构建真正智能的视觉系统。

原文摘要 · Abstract (English)

Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning models but also enhances their ability to interact meaningfully with humans and the environment. Unlike CNN-based conventional vision models, which are designed to identify objects within a specific image, incorporating commonsense knowledge enables models to interpret scenes in a more holistic manner, thereby improving their spatial ability to reason about relationships among objects and actions. This integration not only enhances object recognition but also facilitates a deeper understanding of the contextual factors, ultimately leading to more precise predictions and interactions in real-world applications. This paper presents a comprehensive survey of recent developments that integrate commonsense knowledge into computer vision tasks. We systematically review approaches based on knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. We also outline current limitations related to dataset bias, knowledge incompleteness, and integration challenges. Finally, we highlight prospective research trajectories in cross-modal reasoning, scalable commonsense knowledge injection, and neuro-symbolic hybrid architectures to develop truly intelligent visual systems.

常识推理视觉理解神经符号跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。