仅用头手追踪数据,同时还原人体动作、物体运动与交互关系。
ECHO: Ego-Centric modeling of Human-Object interactions
- 采用三变量扩散模型,联合建模人、物与交互的依赖关系。
- 在多个数据集上表现超越现有方法,尤其在信号中断时更稳定。
- 适合可穿戴设备应用,如智能眼镜和手表场景下的动作分析。
从第一人称视角建模人-物交互是关键但极具挑战的任务,尤其当依赖智能眼镜、手表等可穿戴设备的稀疏信号时。我们提出ECHO,首个仅基于头部与手腕追踪数据,联合恢复人体姿态、物体运动及接触动力学的统一框架。为应对该问题的欠约束特性,引入一种具有独立噪声调度的新型三变量扩散过程,建模人、物与交互模态间的相互依赖。此设计使ECHO可适应灵活输入配置,对追踪间断具有鲁棒性,并能利用大规模人体动作数据集与小规模人-物交互数据集联合训练,既学习强先验又捕捉交互细节。此外,采用平滑插值推理机制,支持任意长序列的时序一致交互生成。大量实验表明,ECHO性能达到当前最优,显著优于缺乏此类灵活性的方法。项目页面见https://ptrvilya.github.io/echo/。
原文摘要 · Abstract (English)
Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified framework to jointly recover human pose, object motion, and contact dynamics solely from head and wrist tracking. To tackle the underconstrained nature of this problem, we introduce a novel tri-variate diffusion process with independent noise schedules that models the mutual dependencies between the human, object, and interaction modalities. This formulation allows ECHO to operate with flexible input configurations, making it robust to intermittent tracking and capable of leveraging partial observations. Crucially, it enables training on a combination of large-scale human motion datasets and smaller HOI collections, learning strong priors while capturing interaction nuances. Furthermore, we employ a smooth inpainting inference mechanism that enables the generation of temporally consistent interactions for arbitrarily long sequences. Extensive evaluations demonstrate that ECHO achieves state-of-the-art performance, significantly outperforming existing methods lacking such flexibility. The project page is available at https://ptrvilya.github.io/echo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。