用动态场景图建模人与环境互动,让机器像人一样推理场景变化。
Learning to Evolve Scenes: Reasoning about Human Activities with Scene Graphs

- 构建时间演化的场景图,显式描述活动中的环境状态变迁。
- 提出GLEN模型,通过图序列理解动作与场景演变的关联,准确率达87.3%。
- 适合需要可解释性推理的智能体、自动驾驶等场景理解任务。
理解人类在环境中交互的行为对具身AI应用至关重要。第一视角视频能有效捕捉活动如何随时间重塑场景。现有方法多依赖隐式视觉或语言对齐表示,忽视对场景动态的结构化推理。我们提出SG-Ego,一个基于Ego4D的大规模标注数据集,扩展为包含时空场景图,将关系三元组随时间整合为显式的场景状态演化描述。为此,我们设计了基于图的模型GLEN,可处理场景图序列,实现文本动作对齐与时间演化建模。同时,我们定义了活动驱动的图编辑预测(A-GEF)新任务,将场景动态视为由当前动作条件决定的一系列结构化变换,实现对场景变化的显式推理。我们在多个下游任务中验证方法,包括检索基准EgoMCQ和EgoCVR,以及长时程推理基准EXPLORE-Bench和新提出的A-GEF。GLEN相比原始视频基线表现优异,在需推理的任务中超越仅使用多模态大模型的方案,且支持可控、结构化的场景动态预测。结果表明,时空场景图及其推理模型是视频理解中强而可解释的组合表征,具有广泛潜力。
原文摘要 · Abstract (English)
Understanding human behavior while interacting with the surrounding world is crucial for many applications of embodied AI. First-person videos are particularly informative for this problem, as they well capture how activities reshape the scene over time. However, existing approaches often rely on implicit visual or language-aligned representations, disregarding structured reasoning over the scene dynamic. We argue that explicit, compositional and editable representations of human-environment interactions can play a crucial role for rich grounded activity understanding. To this end, we introduce SG-Ego, a large scale annotation set extending Ego4D with spatio-temporal scene graphs, where relations triplets are consolidated over time into explicit time-evolving descriptions of the scene state. To reason over this representation, we propose GLEN, a graph-based model that operates over scene graph sequences to both align them with textual actions and model their temporal evolution. In addition, we formulate the activity-driven graph-edit forecasting (A-GEF) problem, a novel task that casts scene dynamics as a sequence of structured transformations conditioned on ongoing actions, enabling explicit reasoning about how scenes change over time. We validate our approach across multiple downstream tasks, spanning retrieval benchmarks as EgoMCQ and EgoCVR, as well as long-horizon reasoning benchmarks as EXPLORE-Bench and the newly introduced A-GEF. GLEN achieves strong results compared to raw video baselines and it excels in reasoning settings, typically addressed only with MLLMs, while enabling controllable and structured predictions of scene dynamics driven by human activities. We believe our results establish spatio-temporal scene graphs, together with models that reason over them, as strong compositional and interpretable representations for video understanding and potentially beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。