arXiv:2603.25726cs.CV2026-03被引 2

用250万张合成图像提升手部姿态估计精度

AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation

  • 构建250万张含深度的合成手部图像数据集
  • 在多个基准上显著提升现有模型性能
  • 适合研究手部姿态与多模态融合的开发者

我们提出AnyHand,一个大规模合成数据集,用于推动3D手部姿态估计的发展。尽管现有基础模型表明数据规模对性能提升至关重要,但真实世界数据集覆盖有限,而以往合成数据集很少同时具备遮挡、手臂细节和对齐深度图。AnyHand包含250万张单手和410万张手物交互的RGB-D图像,附带丰富的几何标注。实验显示,仅通过扩展现有RGB基线的训练数据(不改变模型结构与训练策略),AnyHand即可在FreiHAND和HO-3D等多个基准上带来显著性能提升。结合对训练数据规模与构成的详尽消融分析,结果表明数据多样性与质量与数据规模同样关键。附录还验证了对齐深度图的有效性,表明使用AnyHand扩展RGB-D监督,可使轻量级深度融合模型超越现有RGB-D方法。

原文摘要 · Abstract (English)

We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation. While recent works with foundation approaches have shown that scaling training data markedly improves hand pose estimation, existing real-world datasets are limited in coverage, and prior synthetic datasets rarely provide occlusions, arm details, and aligned depth together at scale. To address this bottleneck, our proposed AnyHand contains 2.5M single-hand and 4.1M hand-object interaction RGB-D images, with rich geometric annotations. We show that extending the original training data recipes of existing RGB baselines with AnyHand yields significant gains on multiple benchmarks (FreiHAND and HO-3D), even when keeping the architectures and training schemes fixed. Together with extensive ablations on the scale and composition of the training data setups, these results suggest that training data diversity and quality are as critical as scale for advancing hand pose estimation. We further examine the utility of AnyHand's aligned depth maps in the appendix, showing that scaling RGB-D supervision with AnyHand allows a lightweight depth-fusion variant of existing RGB baselines to outperform prior RGB-D methods.

手部姿态合成数据深度融合三维重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。