用关键点控制手部生成,实现姿态调整和外观迁移。
FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation

- 以2D关键点为通用表示,统一编码手部动作与视角变化。
- 可零样本修复生成图像中的畸形手,支持视频序列合成。
- 适合需要精细手部控制的生成任务,如虚拟试衣、动画制作。
尽管图像生成模型取得显著进展,但因手部结构复杂、视角多变及频繁遮挡,生成逼真手部仍具挑战。本文提出FoundHand,一种大规模领域特定扩散模型,用于合成单手与双手图像。为训练该模型,我们构建了FoundHand-10M数据集,包含1000万张带2D关键点和分割掩码标注的手部图像。核心思想是利用2D手部关键点作为通用表征,同时编码手部动作与相机视角。FoundHand通过图像对学习,捕捉物理上合理的手部动作,原生支持通过2D关键点精确控制,并具备外观控制能力。模型具备重定姿态、转移手部外观及合成新视角等核心功能,实现零样本修复生成图像中畸形手或合成手部视频序列的能力。大量实验表明,本方法达到当前最优性能。
原文摘要 · Abstract (English)
Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale domain-specific diffusion model for synthesizing single and dual hand images. To train our model, we introduce FoundHand-10M, a large-scale hand dataset with 2D keypoints and segmentation mask annotations. Our insight is to use 2D hand keypoints as a universal representation that encodes both hand articulation and camera viewpoint. FoundHand learns from image pairs to capture physically plausible hand articulations, natively enables precise control through 2D keypoints, and supports appearance control. Our model exhibits core capabilities that include the ability to repose hands, transfer hand appearance, and even synthesize novel views. This leads to zero-shot capabilities for fixing malformed hands in previously generated images, or synthesizing hand video sequences. We present extensive experiments and evaluations that demonstrate state-of-the-art performance of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。