用扩散模型预测人与环境互动的动态变化,不依赖固定物体假设。
FORESCENE: FOREcasting human activity via latent SCENE graphs diffusion
- 用图自编码器将视频转为潜在图表示,再用扩散模型预测未来关系演变。
- 在Action Genome上优于现有方法,能处理物体增减的长期交互任务。
- 适合研究长期行为预测、人环交互建模的科研人员参考。
日常活动中人类与环境的交互预测因行为高度多变而困难。直接从视频预测受限于无关物体或背景噪声等干扰因素。基于场景图(SGs)的方法可仅追踪相关元素,但现有方法常假设物体不变,难以应用于物体可能出现或消失的长期活动。本文提出FORESCENE,一种新型场景图前瞻(SGA)框架,可同时预测物体与关系随时间的演化。该方法通过定制图自编码器将观测视频段编码为潜在表示,并利用潜在扩散模型(LDM)预测未来场景图。该方法无需对图内容或结构做假设,实现连续交互动态预测。我们在Action Genome数据集上评估,结果表明其性能超越现有SGA方法,且解决了更复杂的任务。
原文摘要 · Abstract (English)
Forecasting human-environment interactions in daily activities is challenging due to the high variability of human behavior. While predicting directly from videos is possible, it is limited by confounding factors like irrelevant objects or background noise that do not contribute to the interaction. A promising alternative is using Scene Graphs (SGs) to track only the relevant elements. However, current methods for forecasting future SGs face significant challenges and often rely on unrealistic assumptions, such as fixed objects over time, limiting their applicability to long-term activities where interacted objects may appear or disappear. In this paper, we introduce FORESCENE, a novel framework for Scene Graph Anticipation (SGA) that predicts both object and relationship evolution over time. FORESCENE encodes observed video segments into a latent representation using a tailored Graph Auto-Encoder and forecasts future SGs using a Latent Diffusion Model (LDM). Our approach enables continuous prediction of interaction dynamics without making assumptions on the graph's content or structure. We evaluate FORESCENE on the Action Genome dataset, where it outperforms existing SGA methods while solving a significantly more complex task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。