通过显式骨骼推理,让生成的人像更自然、结构更合理。
SkeleGuide: Explicit Skeleton Reasoning for Context-Aware Human-in-Place Image Synthesis
- 引入骨骼推理模块,用内部姿态引导图像生成
- 在多个数据集上优于主流模型,姿态更自然
- 支持用户编辑姿态,适合需要精准控制的场景
将真实人体图像合成到现有场景中仍面临巨大挑战,现有生成模型常出现肢体扭曲和姿态异常。我们归因于缺乏对人类骨骼结构的显式推理能力。为此,提出SkeleGuide框架,通过联合训练推理与渲染阶段,学习生成一个内部姿态作为强结构先验,指导合成过程保持高结构一致性。为实现细粒度控制,引入PoseInverter模块,可将内部隐式姿态解码为显式可编辑格式。大量实验表明,SkeleGuide在生成高质量、上下文感知的人体图像方面显著优于专用与通用模型。本工作证明,显式建模骨骼结构是实现稳健且逼真人像合成的关键一步。
原文摘要 · Abstract (English)
Generating realistic and structurally plausible human images into existing scenes remains a significant challenge for current generative models, which often produce artifacts like distorted limbs and unnatural poses. We attribute this systemic failure to an inability to perform explicit reasoning over human skeletal structure. To address this, we introduce SkeleGuide, a novel framework built upon explicit skeletal reasoning. Through joint training of its reasoning and rendering stages, SkeleGuide learns to produce an internal pose that acts as a strong structural prior, guiding the synthesis towards high structural integrity. For fine-grained user control, we introduce PoseInverter, a module that decodes this internal latent pose into an explicit and editable format. Extensive experiments demonstrate that SkeleGuide significantly outperforms both specialized and general-purpose models in generating high-fidelity, contextually-aware human images. Our work provides compelling evidence that explicitly modeling skeletal structure is a fundamental step towards robust and plausible human image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。