让手控视频生成摆脱实验室限制,实现真实场景下的自然控制。
HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control

- 用主角中心标注法从单目视频中提取高质量手部轨迹
- 构建103K片段的无约束手部数据集,支持复杂相机运动
- 提出普吕克手部映射,分离相机与手部运动,提升控制精度
我们提出HandsOnWorld,一个直接从无约束单目视频中学习的手控第一人称视频生成框架。现有生成方法依赖多视角或标记式动作捕捉提供的3D手部标注,仅限于受控实验环境。为弥补这一差距,我们设计了一种主角中心标注流程,在动作语义、图像质量与3D几何层面过滤单目3D重建结果,构建了包含103,000个片段、约1200万帧的EgoVid-Pro数据集,覆盖多样日常场景。这些无约束场景中存在显著相机自运动,而传统相机空间控制信号常将相机与手部运动混杂。为此,我们提出普吕克手部映射(Plücker Hand Map),将普吕克射线从相机几何扩展至手部表面,以统一世界坐标系表示手部运动,从表征层面解耦两类运动源。实验表明,HandsOnWorld在视觉保真度和控制准确性上优于现有方法,并可泛化至实验室外真实场景。
原文摘要 · Abstract (English)
We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand annotations from multi-view or marker-based motion capture, confining them to narrow, instrumented scene distributions. To bridge this gap, we introduce a protagonist-centered annotation pipeline that filters monocular 3D reconstructions at the action-semantic, image-quality, and 3D-geometric levels, yielding EgoVid-Pro, a dataset of clean, protagonist-only hand trajectories spanning 103K clips and roughly 12M frames across diverse everyday scenes. These unconstrained scenes exhibit substantial camera ego-motion that is largely absent from tabletop captures, exposing the entanglement of camera and hand motion in existing camera-space control signals. We therefore propose the Plücker Hand Map, which extends Plücker rays from camera geometry to the hand surface, representing hand motion in the same world frame as the camera and disentangling the two motion sources at the representation level. Experiments show that HandsOnWorld outperforms prior methods in visual fidelity and control accuracy and generalizes beyond laboratory settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。