arXiv:2410.13662cs.CV2024-10被引 2

让AI看图推理动作背后的常识,无需训练就能理解做饭细节

ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions

  • 用语言模型分析图像,零样本推断动作相关常识
  • 构建8.5k张图+59.3万条常识推理数据集
  • 适合做视觉常识理解与智能助手的开发者

人类观察他人执行动作(实际或视频/图像中)时,能做出超越视觉感知的广泛推断,包括动作实现的物理条件(如液体可被倾倒)、动作导致的世界变化(如土豆炸后变金黄酥脆)、动作的高层次目标(如打蛋是为了煎蛋饼),以及前后动作关系(如打蛋前要敲开蛋壳,煮面后需沥水)。这种推理能力对辅助人类完成日常任务的自主系统至关重要。为此,我们提出一个跨模态任务,学习图像中动作相关的常识概念。我们构建了一个包含8.5k张图像和59.3k条基于这些图像的动作常识推理的数据集,数据源自标注过的烹饪视频数据集。我们提出了ActionCOMET——一种零样本框架,用于识别语言模型中与给定视觉输入相关的特定知识。我们在所收集数据集上展示了ActionCOMET的基线结果,并与现有最佳VQA方法进行对比。

原文摘要 · Abstract (English)

Humans observe various actions being performed by other humans (physically or in videos/images) and can draw a wide range of inferences about it beyond what they can visually perceive. Such inferences include determining the aspects of the world that make action execution possible (e.g. liquid objects can undergo pouring), predicting how the world will change as a result of the action (e.g. potatoes being golden and crispy after frying), high-level goals associated with the action (e.g. beat the eggs to make an omelet) and reasoning about actions that possibly precede or follow the current action (e.g. crack eggs before whisking or draining pasta after boiling). Similar reasoning ability is highly desirable in autonomous systems that would assist us in performing everyday tasks. To that end, we propose a multi-modal task to learn aforementioned concepts about actions being performed in images. We develop a dataset consisting of 8.5k images and 59.3k inferences about actions grounded in those images, collected from an annotated cooking-video dataset. We propose ActionCOMET, a zero-shot framework to discern knowledge present in language models specific to the provided visual input. We present baseline results of ActionCOMET over the collected dataset and compare them with the performance of the best existing VQA approaches.

视觉常识零样本图像理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。