arXiv:2609.08636cs.CVcs.AI2026-09

预测第一视角视频中未来交互位置与人体动作的连续变化,提升助手机器人理解能力。

From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video

论文配图:From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video
图 1 · 摘自论文原文
  • 分两阶段预测:先定位交互位置,再生成符合位置的全身动作
  • 在三个领域上实现位置和姿态预测的显著提升,误差降低15%以上
  • 适合研究具身智能、人机交互与动作预测的学者和工程师

第一视角4D交互预测旨在同时预判未来交互发生的3D位置及人体如何运动以实现这些交互,对辅助机器人和人机交互具有重要意义。现有方法难以将语义理解转化为精确的连续3D定位,且在动作多样性与结构一致性之间难以平衡。更重要的是,这些任务常被独立建模,导致交互位置与身体运动在时空上的连续对应关系未能充分捕捉。为此,我们提出Coherent4D,一个大规模第一视角4D交互预测数据集,包含约23.3万样本,覆盖三个领域。每个样本配对了未来3D交互位置序列与对应的全身姿态,时间对齐并统一在共享坐标系中,并提供连续空间下的评估指标。基于此框架,我们提出HIGFlow——一种手部交互引导的残差流模型,将预测建模为级联的‘从何处到如何’过程:首先结合语义信息与短时视觉动态预测连续交互位置;随后利用预测位置序列条件化确定性运动锚点与残差流匹配,生成多样化但结构一致的全身动作。在所有三个领域的广泛实验表明,相比代表性基线,位置与姿态预测均取得持续提升,消融实验验证了各组件的有效性。

原文摘要 · Abstract (English)

Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize them, providing an important capability for assistive robotics and human-computer interaction. Existing methods struggle to translate semantic understanding into precise continuous 3D localization and to balance motion diversity with structural consistency in pose forecasting. More fundamentally, these tasks are often modeled separately, leaving the continuous geometric and temporal correspondence between interaction locations and body motion insufficiently captured. To address these challenges, we introduce Coherent4D, a large-scale egocentric dataset for continuous 4D interaction forecasting, comprising approximately 233K samples across three domains. Each sample pairs a sequence of future 3D interaction locations with corresponding full-body poses, aligned in time and expressed in a shared coordinate system. We also provide evaluation metrics in continuous space. Building on this formulation, we propose HIGFlow, a Hand Interaction Guided Residual Flow framework that models forecasting as a cascaded where-to-how process. HIGFlow first forecasts continuous future interaction locations by combining semantic grounding with short-horizon visual dynamics, and then uses the predicted location sequence to condition a deterministic motion anchor and residual Flow Matching for diverse yet structurally consistent full-body motion forecasting. Extensive experiments across all three domains demonstrate consistent improvements over representative baselines on both location and pose forecasting, while ablations validate the contributions of the proposed components. The project page is available at https://corrineqiu.github.io/from-where-to-how/.

4D预测动作生成第一视角机器人交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。