arXiv:2606.28215cs.CVcs.AI2026-06中稿 · ECCV

用单视角视频重建多物体4D动态交互,提升智能体训练数据质量

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

论文配图:HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
图 1 · 摘自论文原文
  • 融合视觉语言模型与人机协作反馈,解决深度模糊和遮挡问题
  • 在MVOIK-4D基准上达到当前最优性能,物理合理性显著提升
  • 适合需要真实世界多物体交互数据的机器人与虚拟智能体研究者

从海量真实场景的单目视频中提取动态4D物体交互,为具身AI扩展和视觉语言模型训练提供高效的数据采集路径。然而,现有单目4D重建方法多聚焦于孤立物体,在多物体交互带来的严重遮挡与复杂动态下表现不佳。为此,我们提出HAT-4D,首个专为从单个视频中重建多物体3D几何、时间动态及物理交互而设计的智能体框架。通过将视觉语言模型与多层级人机协同反馈机制结合,HAT-4D在3D生成与4D演化过程中有效化解深度歧义与交互导致的遮挡,无需昂贵多摄像机设备即可生成物理合理资产。作为可扩展的数据引擎,HAT-4D构建了开放世界基准MVOIK-4D,并引入新的多维度评估协议,重点衡量物理合理性与时序一致性。大量实验表明,HAT-4D在多数指标上达到当前最优,同时保持良好的语义对齐。消融实验证明,少量人工反馈显著提升交互重建效果。此外,使用HAT-4D生成的数据微调基线模型后,性能明显提升。代码与数据已开源。

原文摘要 · Abstract (English)

Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs. However, existing monocular 4D reconstruction methods primarily focus on isolated objects, often failing under the severe occlusions and complex dynamics inherent in multi-object interactions. To bridge this gap, we propose HAT-4D, the first agentic framework designed to reconstruct the 3D geometry, temporal dynamics, and physical interactions of multiple objects from a single video. By integrating VLMs with a multi-level human-in-the-loop feedback mechanism, HAT-4D efficiently resolves depth ambiguities and interaction-induced occlusions during 3D generation and 4D propagation, yielding physically plausible assets without relying on expensive multicamera rigs. As a scalable data engine, HAT-4D facilitates the creation of MVOIK-4D, an open-world benchmark for monocular 4D interaction reconstruction, accompanied by a novel multi-dimensional evaluation protocol focused on physical plausibility and temporal consistency. Extensive experiments demonstrate that HAT-4D achieves SOTA performance on most evaluation metrics, while maintaining competitive semantic alignment. Ablation studies show that introducing a small amount of human feedback improves interaction reconstruction. Moreover, the data produced by HAT-4D effectively improves baseline performance when used for fine-tuning. Our data and code are available at https://lijiaxin0111.github.io/HAT4D/

4D重建人机协作多物体交互具身AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。