arXiv:2508.01835cs.CV2025-08ICCV被引 4

用物理规律指导扩散模型,提升单目视频中手部动作的精准重建。

Diffusion-based 3D Hand Motion Recovery with Intuitive Physics

  • 基于扩散模型迭代去噪,结合物理约束优化手部运动估计。
  • 仅用动捕数据训练,在多个基准上达到当前最佳效果。
  • 特别适合手物交互场景,提升动作连贯性与真实性。

尽管从单目图像进行3D手部重建已取得显著进展,但从视频中生成准确且时间连贯的动作估计仍具挑战性,尤其是在手物交互过程中。本文提出一种新型3D手部运动恢复框架,通过基于扩散模型并融合物理规律的运动精修机制,增强图像驱动的重建结果。该模型在初始估计基础上,学习精修后运动序列的分布,通过迭代去噪过程生成更优序列。训练不依赖稀缺的带标注视频数据,仅使用无图像的动捕数据。我们识别出手物交互中的关键运动状态及其对应运动约束等直观物理知识,并有效融入扩散模型以提升性能。大量实验表明,该方法显著改进多种帧级重建方法,在现有基准上实现顶尖表现。

原文摘要 · Abstract (English)

While 3D hand reconstruction from monocular images has made significant progress, generating accurate and temporally coherent motion estimates from videos remains challenging, particularly during hand-object interactions. In this paper, we present a novel 3D hand motion recovery framework that enhances image-based reconstructions through a diffusion-based and physics-augmented motion refinement model. Our model captures the distribution of refined motion estimates conditioned on initial ones, generating improved sequences through an iterative denoising process. Instead of relying on scarce annotated video data, we train our model only using motion capture data without images. We identify valuable intuitive physics knowledge during hand-object interactions, including key motion states and their associated motion constraints. We effectively integrate these physical insights into our diffusion model to improve its performance. Extensive experiments demonstrate that our approach significantly improves various frame-wise reconstruction methods, achieving state-of-the-art (SOTA) performance on existing benchmarks.

3D重建扩散模型物理引导手部动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。