联合建模相机与人体运动,提升单目视频中全局姿态重建精度
WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
- 提出解析式朝向分解法,高效解耦相机与人体运动
- 设计相机轨迹融合机制,显著提升世界坐标系下轨迹重建效果
- 适合做虚拟现实、机器人动作捕捉的开发者参考
从自然场景单目视频中进行全局人体动作重建在虚拟现实、图形学和机器人领域需求日益增长,但需将人体姿态从相机坐标准确映射到世界坐标,该任务受深度模糊、运动模糊以及相机与人体运动纠缠的挑战。现有以人体动作为中心的方法虽能保持动作细节和物理合理性,却存在两个关键缺陷:未能充分挖掘相机朝向信息,且对相机平移线索整合效率低。本文提出WATCH(World-aware Allied Trajectory and pose reconstruction for Camera and Human),一个统一框架解决上述问题。方法引入解析式航向角分解技术,相比现有几何方法更具效率和可扩展性;同时设计受世界模型启发的相机轨迹融合机制,有效利用相机平移信息,超越传统硬解码方式。在真实场景基准测试中,WATCH实现端到端轨迹重建的最先进性能。本工作验证了联合建模相机-人体运动关系的有效性,为长期存在的相机平移整合难题提供新思路。代码将公开。
原文摘要 · Abstract (English)
Global human motion reconstruction from in-the-wild monocular videos is increasingly demanded across VR, graphics, and robotics applications, yet requires accurate mapping of human poses from camera to world coordinates-a task challenged by depth ambiguity, motion ambiguity, and the entanglement between camera and human movements. While human-motion-centric approaches excel in preserving motion details and physical plausibility, they suffer from two critical limitations: insufficient exploitation of camera orientation information and ineffective integration of camera translation cues. We present WATCH (World-aware Allied Trajectory and pose reconstruction for Camera and Human), a unified framework addressing both challenges. Our approach introduces an analytical heading angle decomposition technique that offers superior efficiency and extensibility compared to existing geometric methods. Additionally, we design a camera trajectory integration mechanism inspired by world models, providing an effective pathway for leveraging camera translation information beyond naive hard-decoding approaches. Through experiments on in-the-wild benchmarks, WATCH achieves state-of-the-art performance in end-to-end trajectory reconstruction. Our work demonstrates the effectiveness of jointly modeling camera-human motion relationships and offers new insights for addressing the long-standing challenge of camera translation integration in global human motion reconstruction. The code will be available publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。