arXiv:2506.08553cs.CV2025-06被引 5

用场景图与常识图提升第一人称视觉问答能力

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

  • 通过多模态大模型生成场景图,捕捉物体交互与时空关系
  • 融合ConceptNet常识知识,实现超越视觉证据的推理
  • 在7类任务中表现均衡,综合准确率达44.21%

本报告介绍为应对HD-EPIC VQA Challenge 2025提出的SceneNet与KnowledgeNet方法。SceneNet利用多模态大语言模型(MLLM)生成的场景图,捕捉细粒度的物体交互、空间关系及时间定位事件;同时,KnowledgeNet引入ConceptNet的外部常识知识,建立实体间的高层语义关联,实现对直接视觉证据之外的推理。两种方法在HD-EPIC基准的七个类别中各具优势,其组合框架在挑战赛中取得44.21%的整体准确率,验证了其在复杂第一人称视觉问答任务中的有效性。

原文摘要 · Abstract (English)

This report presents SceneNet and KnowledgeNet, our approaches developed for the HD-EPIC VQA Challenge 2025. SceneNet leverages scene graphs generated with a multi-modal large language model (MLLM) to capture fine-grained object interactions, spatial relationships, and temporally grounded events. In parallel, KnowledgeNet incorporates ConceptNet's external commonsense knowledge to introduce high-level semantic connections between entities, enabling reasoning beyond directly observable visual evidence. Each method demonstrates distinct strengths across the seven categories of the HD-EPIC benchmark, and their combination within our framework results in an overall accuracy of 44.21% on the challenge, highlighting its effectiveness for complex egocentric VQA tasks.

视觉问答场景图常识推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。