通过模拟局部解剖错误,提升文本生成人像的结构准确性。
Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences

- 用可控退化机制在高质量人像中引入局部解剖错误,构建偏好数据对。
- 在10,000+数据对上训练,显著减少生成图像的解剖错误。
- 专为评估人体结构真实性设计基准测试,适合关注生成质量的研究者。
大规模文本到图像基础模型虽已实现卓越的视觉真实感,但在生成具有正确解剖结构的人像方面仍面临挑战。现有方法通过特定部位模块或局部损失加权在高质量人像数据上进行监督微调,但此类数据有限,且受光照、姿态、背景等混淆因素影响,优化信号模糊。偏好对齐提供替代方案,但标准直接偏好优化(DPO)对所有像素一视同仁,无法利用解剖缺陷的局部特性。为此,我们提出基于合成解剖偏好对的对齐框架(ASAP),通过在高保真人像中施加局部退化机制,构造受控偏好对。该机制在特定区域引入明确的解剖错误,同时保留其余内容。基于此,我们构建了包含超10,000对的「人类解剖偏好」(HAP)数据集,用于有效对齐文本到图像人像生成模型。为更好利用这些受控偏好对的局部性,我们引入一种局部且带边界限制的DPO变体,优先在目标解剖区域优化,同时设置有限偏好边界以防止过度优化并保留全局语义。我们进一步提出了HAF-Bench,一个系统评估解剖保真度的基准。大量实验表明,ASAP在多个基础模型上持续降低解剖错误,同时保持整体图像质量。
原文摘要 · Abstract (English)
Large-scale text-to-image foundation models have achieved remarkable visual realism, yet generating human images with correct anatomical structures remains challenging. Existing approaches enforce anatomical constraints through part-specific modules or localized loss weighting during supervised fine-tuning on high-quality human photos, but such datasets are limited and often provide ambiguous optimization signals due to confounding factors such as lighting, pose, and background. Preference-based alignment offers an alternative, but standard Direct Preference Optimization (DPO) treats all pixels equally and therefore fails to exploit the localized nature of anatomical artifacts. To address this, we propose the framework of Alignment via Synthetic Anatomical Preference (ASAP), which constructs controlled preference pairs through a localized degradation mechanism applied to high-fidelity human images. This mechanism performs a controlled experiment on images by introducing explicit anatomical errors in targeted regions while preserving the remaining content. With this mechanism, we create the Human Anatomical Preference (HAP) dataset with over 10K curated pairs for effective anatomical alignment of text-to-image human image generative models. To better leverage the locality of these controlled preference pairs, we introduce a localized and margin-bounded variant of DPO that prioritizes optimization in targeted anatomical regions while enforcing a finite preference margin to prevent over-optimization and preserve global semantics. We further introduce HAF-Bench, a benchmark for systematic evaluation of anatomical fidelity. Extensive experiments demonstrate that ASAP consistently reduces anatomical errors across multiple foundation models while maintaining overall image quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。