提出情境场景图,让机器更懂人类行为的细节环境。
Situational Scene Graph for Structured Human-centric Situation Understanding
- 用预定义角色值建模动作的细粒度语义属性
- 新数据集标注了人、物、动词的语义角色-值对
- 统一表示提升关系分类与情境推理性能
基于图的表示方法广泛用于视频理解中的时空关系建模。尽管有效,现有方法多关注人-物关系,忽视动作组件的细粒度语义属性,如动作发生地点、使用工具及物体功能属性。这些属性对理解当前情境至关重要。本文提出一种名为情境场景图(Situational Scene Graph, SSG)的图结构表示,同时编码人-物关系及其对应的语义属性。语义信息以情境框架中预定义的角色与值形式呈现,原用于单个动作表征。基于该表示,我们引入情境场景图生成任务,并提出多阶段交互互补网络(InComNet)解决该任务。由于现有数据集不适用于此任务,我们构建了一个新的SSG数据集,其标注包含人、物及人-物关系谓词的语义角色-值框架。实验表明,该统一表示不仅能提升谓词分类与语义角色-值分类性能,还能增强以人为核心的场景理解推理能力。代码与数据集将陆续公开。
原文摘要 · Abstract (English)
Graph based representation has been widely used in modelling spatio-temporal relationships in video understanding. Although effective, existing graph-based approaches focus on capturing the human-object relationships while ignoring fine-grained semantic properties of the action components. These semantic properties are crucial for understanding the current situation, such as where does the action takes place, what tools are used and functional properties of the objects. In this work, we propose a graph-based representation called Situational Scene Graph (SSG) to encode both human-object relationships and the corresponding semantic properties. The semantic details are represented as predefined roles and values inspired by situation frame, which is originally designed to represent a single action. Based on our proposed representation, we introduce the task of situational scene graph generation and propose a multi-stage pipeline Interactive and Complementary Network (InComNet) to address the task. Given that the existing datasets are not applicable to the task, we further introduce a SSG dataset whose annotations consist of semantic role-value frames for human, objects and verb predicates of human-object relations. Finally, we demonstrate the effectiveness of our proposed SSG representation by testing on different downstream tasks. Experimental results show that the unified representation can not only benefit predicate classification and semantic role-value classification, but also benefit reasoning tasks on human-centric situation understanding. We will release the code and the dataset soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。