arXiv:2501.09041cs.CVcs.CL2025-01被引 1

用生成的场景图提升视觉常识推理的准确性和解释性。

Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing

  • 先用图像块和大模型生成无位置场景图,再基于图推理答案。
  • 在VQA-CC和SG-CR数据集上,答案准确率提升7.3%以上。
  • 适合需要可解释视觉推理的研究者和开发者。

视觉常识推理被视为衡量AI系统高级场景理解能力的关键挑战任务。然而,可靠的推理依赖于对场景细节的充分掌握。现有方法未能有效利用场景中真实存在的物体关系信息,过度依赖训练记忆中的知识。为此,我们提出一种基于场景图增强的视觉常识推理生成方法G2:首先利用图像块和大语言模型生成无位置的场景图,再基于场景图信息进行答案生成与解释。同时,设计了自动场景图过滤与选择策略,在训练中吸收有价值的图结构信息。在场景图构建和视觉常识问答与解释任务上进行了大量实验。结果与消融分析表明,所提框架具有显著有效性。

原文摘要 · Abstract (English)

Visual Commonsense Reasoning, which is regarded as one challenging task to pursue advanced visual scene comprehension, has been used to diagnose the reasoning ability of AI systems. However, reliable reasoning requires a good grasp of the scene's details. Existing work fails to effectively exploit the real-world object relationship information present within the scene, and instead overly relies on knowledge from training memory. Based on these observations, we propose a novel scene-graph-enhanced visual commonsense reasoning generation method named \textit{\textbf{G2}}, which first utilizes the image patches and LLMs to construct a location-free scene graph, and then answer and explain based on the scene graph's information. We also propose automatic scene graph filtering and selection strategies to absorb valuable scene graph information during training. Extensive experiments are conducted on the tasks and datasets of scene graph constructing and visual commonsense answering and explaining, respectively. Experimental results and ablation analysis demonstrate the effectiveness of our proposed framework.

视觉推理场景图生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。