提出通用手部动作预测框架,支持多目标、多维度、多任务预测。
Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views

- 融合视觉语言与上下文信息,实现跨模态统一建模。
- 首次在真实任务中评估,性能超越现有方法。
- 适合机器人操控与动作预判场景,提升实际应用能力。
在第一人称视角下预测手部运动对增强现实和人机策略迁移至关重要。现有手部轨迹预测方法存在目标不足、模态差异、手头动作耦合及下游任务验证有限等问题。为此,本文提出统一手部动作预测框架,支持多模态输入、多维多目标预测及多任务泛化。通过视觉-语言融合、全局上下文建模和任务感知文本嵌入注入,实现2D/3D空间中手部关键点的联合预测。设计新型双分支扩散模型,同步预测头部与手部运动,捕捉其协同关系。引入目标指示符,可精准预测腕关节或手指关键点,而不仅限于手中心点。此外,还新增手物交互状态(接触/分离)预测,更好支持下游任务。作为首个在真实任务中评估的方法,我们构建新基准测试。在多个公开数据集及自建基准上,Uni-Hand 在多维多目标预测中达到领先性能;在机器人策略迁移与动作预判等任务中表现优异。
原文摘要 · Abstract (English)
Forecasting how human hands move in egocentric views is critical for applications like augmented reality and human-robot policy transfer. Recently, several hand trajectory prediction (HTP) methods have been developed to generate future possible hand waypoints, which still suffer from insufficient prediction targets, inherent modality gaps, entangled hand-head motion, and limited validation in downstream tasks. To address these limitations, we present a universal hand motion forecasting framework considering multi-modal input, multi-dimensional and multi-target prediction patterns, and multi-task affordances for downstream applications. We harmonize multiple modalities by vision-language fusion, global context incorporation, and task-aware text embedding injection, to forecast hand waypoints in both 2D and 3D spaces. A novel dual-branch diffusion is proposed to concurrently predict human head and hand movements, capturing their motion synergy in egocentric vision. By introducing target indicators, the prediction model can forecast the specific joint waypoints of the wrist or the fingers, besides the widely studied hand center points. In addition, we enable Uni-Hand to additionally predict hand-object interaction states (contact/separation) to facilitate downstream tasks better. As the first work to incorporate downstream task evaluation in the literature, we build novel benchmarks to assess the real-world applicability of hand motion forecasting algorithms. The experimental results on multiple publicly available datasets and our newly proposed benchmarks demonstrate that Uni-Hand achieves the state-of-the-art performance in multi-dimensional and multi-target hand motion forecasting. Extensive validation in multiple downstream tasks also presents its impressive human-robot policy transfer to enable robotic manipulation, and effective feature enhancement for action anticipation/recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。