预测视频中人未来与物体的4D交互,包含位置和动作。
FIction: 4D Future Interaction Prediction from Video
- 融合视频时序信息,预测未来交互的3D位置与动作姿态。
- 在EgoExo4D数据集上相对基线提升超30%。
- 适合需要精准动作预判的机器人与AR应用。
预判人如何与环境中的物体交互是理解行为的关键,但现有方法仅限于视频帧的2D空间,只能预测“做什么”,忽略“在哪”和“如何做”。本文提出FIction模型,实现从视频中预测4D未来交互。给定一段人类活动视频,目标是预测未来时间段内人将在哪些3D位置与哪些物体互动(如橱柜、冰箱),以及具体执行的动作(如弯腰、伸手、拉拽)。FIction通过融合过去视频中人的动作与环境信息,同时预测交互的“位置”和“方式”。在EgoExo4D的真实场景多类活动中进行充分实验,结果表明该方法显著优于以往自回归及(升维)2D视频模型,相对性能提升超过30%。
原文摘要 · Abstract (English)
Anticipating how a person will interact with objects in an environment is essential for activity understanding, but existing methods are limited to the 2D space of video frames-capturing physically ungrounded predictions of "what" and ignoring the "where" and "how". We introduce FIction for 4D future interaction prediction from videos. Given an input video of a human activity, the goal is to predict which objects at what 3D locations the person will interact with in the next time period (e.g., cabinet, fridge), and how they will execute that interaction (e.g., poses for bending, reaching, pulling). Our novel model FIction fuses the past video observation of the person's actions and their environment to predict both the "where" and "how" of future interactions. Through comprehensive experiments on a variety of activities and real-world environments in EgoExo4D, we show that our proposed approach outperforms prior autoregressive and (lifted) 2D video models substantially, with more than 30% relative gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。