用语义与结构对齐提升3D手部重建的生成质量
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
- 通过视觉语言模型提取行为语义,聚焦手部上下文
- 引入分层结构融合,增强生成图像中手与身体对齐
- 适合需要高质量手部生成数据的研究者
近期3D手部重建研究证明了合成数据在提升估计性能上的有效性。然而,多数方法依赖游戏引擎生成手部图像,往往缺乏纹理与环境多样性,且缺少手臂或交互物体等关键成分。生成模型是替代方案,但仍存在对齐问题。本文提出SesaHand,从语义和结构两个角度增强可控手部图像生成。针对语义对齐,我们设计基于思维链推理的流水线,从视觉语言模型生成的图像描述中提取人类行为语义,抑制非人体细节,确保充分的人体中心上下文。针对结构对齐,提出分层结构融合机制,整合多粒度结构信息以优化特征,更好对齐生成图像中的手与整体人体。此外,提出手部结构注意力增强方法,有效提升模型对关键区域的关注。实验表明,该方法不仅在生成性能上优于现有工作,还显著提升了使用生成图像进行3D手部重建的效果。
原文摘要 · Abstract (English)
Recent studies on 3D hand reconstruction have demonstrated the effectiveness of synthetic training data to improve estimation performance. However, most methods rely on game engines to synthesize hand images, which often lack diversity in textures and environments, and fail to include crucial components like arms or interacting objects. Generative models are promising alternatives to generate diverse hand images, but still suffer from misalignment issues. In this paper, we present SesaHand, which enhances controllable hand image generation from both semantic and structural alignment perspectives for 3D hand reconstruction. Specifically, for semantic alignment, we propose a pipeline with Chain-of-Thought inference to extract human behavior semantics from image captions generated by the Vision-Language Model. This semantics suppresses human-irrelevant environmental details and ensures sufficient human-centric contexts for hand image generation. For structural alignment, we introduce hierarchical structural fusion to integrate structural information with different granularity for feature refinement to better align the hand and the overall human body in generated images. We further propose a hand structure attention enhancement method to efficiently enhance the model's attention on hand regions. Experiments demonstrate that our method not only outperforms prior work in generation performance but also improves 3D hand reconstruction with the generated hand images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。