用梯度引导生成真实物理动作,提升第一人称视角物体运动预测能力
EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation
- 混合Mamba-Transformer-Perceiver架构融合时序、几何与语义信息
- 梯度引导推理降低79%碰撞率,生成更平滑、真实的6自由度轨迹
- 无需后处理即可实现可控生成,适合真实场景下机器人交互应用
理解并预测第一人称视频中的物体运动是具身感知与交互的基础。然而,由于遮挡、快速运动以及现有生成模型缺乏显式物理推理,生成物理一致的6自由度轨迹仍具挑战。我们提出EgoFlow,一种基于流匹配的框架,可从多模态第一人称观测中合成真实且物理合理的运动轨迹。EgoFlow采用混合Mamba-Transformer-Perceiver架构,联合建模时间动态、场景几何与语义意图;同时通过梯度引导推理过程施加可微分的物理约束,如避障与运动平滑性。该方法在真实数据集HD-EPIC、EgoExo4D和HOT3D上表现优异,相比扩散模型与Transformer基线,在准确率、泛化能力与物理真实性方面均有提升,碰撞率最高降低79%,并展现出对未见场景的强大泛化能力。结果表明,基于流的生成建模在可扩展且物理可信的第一人称运动理解中具有巨大潜力。
原文摘要 · Abstract (English)
Understanding and predicting object motion from egocentric video is fundamental to embodied perception and interaction. However, generating physically consistent 6DoF trajectories remains challenging due to occlusions, fast motion, and the lack of explicit physical reasoning in existing generative models. We present EgoFlow, a flow-matching framework that synthesizes realistic and physically plausible trajectories conditioned on multimodal egocentric observations. EgoFlow employs a hybrid Mamba-Transformer-Perceiver architecture to jointly model temporal dynamics, scene geometry, and semantic intent, while a gradient-guided inference process enforces differentiable physical constraints such as collision avoidance and motion smoothness. This combination yields coherent and controllable motion generation without post-hoc filtering or additional supervision. Experiments on real-world datasets HD-EPIC, EgoExo4D, and HOT3D show that EgoFlow outperforms diffusion-based and transformer baselines in accuracy, generalization, and physical realism, reducing collision rates by up to 79%, and strong generalization to unseen scenes. Our results highlight the promise of flow-based generative modeling for scalable and physically grounded egocentric motion understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。