arXiv:2510.06145cs.CVcs.AI2025-10被引 1

从日常图像中预测双手3D运动与关节变化,提升零样本泛化能力

Bimanual 3D Hand Motion and Articulation Forecasting in Everyday Images

  • 用扩散模型将2D关键点序列升维为4D手部运动数据
  • 在6个数据集上实现14%性能提升,零样本测试精度提高16.4%
  • 适合做手势预测、人机交互的开发者和研究者

我们针对从单张日常图像中预测双手3D运动与关节状态的问题展开研究。为解决多样化场景下缺乏3D手部标注的问题,设计了一套注释流程:利用扩散模型将2D手部关键点序列映射为4D手部运动。针对预测模型,采用扩散损失以捕捉手部运动分布的多模态特性。在6个数据集上的大量实验表明,使用插补标签进行训练可带来14%的性能提升,所提出的升维模型表现优于基线42%,预测模型也有16.4%的增益。尤其在零样本泛化至日常图像时,效果显著。

原文摘要 · Abstract (English)

We tackle the problem of forecasting bimanual 3D hand motion & articulation from a single image in everyday settings. To address the lack of 3D hand annotations in diverse settings, we design an annotation pipeline consisting of a diffusion model to lift 2D hand keypoint sequences to 4D hand motion. For the forecasting model, we adopt a diffusion loss to account for the multimodality in hand motion distribution. Extensive experiments across 6 datasets show the benefits of training on diverse data with imputed labels (14% improvement) and effectiveness of our lifting (42% better) & forecasting (16.4% gain) models, over the best baselines, especially in zero-shot generalization to everyday images.

3D手势预测扩散模型多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。