arXiv:2512.22272cs.CV2025-12

用轻量级模型引导扩散模型,让生成图像更符合真实几何形状。

Human-Aligned Generative Perception: Bridging Psychophysics and Generative Models

  • 用人类感知数据训练小模型作为教师,提供几何约束信号。
  • 相比无引导基线,语义对齐提升约80%。
  • 无需重新训练,即可零样本迁移复杂3D形状到冲突材质上。

文本到图像的扩散模型虽能生成高细节纹理,但常依赖表面外观,难以遵循严格的几何约束,尤其当文本提示暗示的风格与几何要求冲突时。这反映了人类感知与当前生成模型间的语义鸿沟。本文探索是否可通过轻量级、现成的判别器作为外部引导信号,在不进行专门训练的前提下引入几何理解。我们提出了基于THINGS三元组数据集训练的人类感知嵌入(HPE)教师模型,捕捉人类对物体形状的敏感度。通过将该教师模型的梯度注入潜在扩散过程,实现了几何与风格的可控分离。在三种架构上验证:以U-Net为骨干的Stable Diffusion v1.5、流匹配模型SiT-XL/2,以及扩散变压器PixArt-Σ。实验表明,流模型在缺乏持续引导时易偏离默认轨迹;我们还展示了复杂三维形状(如Eames椅)可零样本迁移至矛盾材质(如粉红色金属)上。该引导生成方式使语义对齐相比无引导基线提升约80%。总体结果表明,小型教师模型可可靠引导大型生成系统,实现更强几何控制并拓展文本到图像合成的创作空间。

原文摘要 · Abstract (English)

Text-to-image diffusion models generate highly detailed textures, yet they often rely on surface appearance and fail to follow strict geometric constraints, particularly when those constraints conflict with the style implied by the text prompt. This reflects a broader semantic gap between human perception and current generative models. We investigate whether geometric understanding can be introduced without specialized training by using lightweight, off-the-shelf discriminators as external guidance signals. We propose a Human Perception Embedding (HPE) teacher trained on the THINGS triplet dataset, which captures human sensitivity to object shape. By injecting gradients from this teacher into the latent diffusion process, we show that geometry and style can be separated in a controllable manner. We evaluate this approach across three architectures: Stable Diffusion v1.5 with a U-Net backbone, the flow-matching model SiT-XL/2, and the diffusion transformer PixArt-Σ. Our experiments reveal that flow models tend to drift back toward their default trajectories without continuous guidance, and we demonstrate zero-shot transfer of complex three-dimensional shapes, such as an Eames chair, onto conflicting materials such as pink metal. This guided generation improves semantic alignment by about 80 percent compared to unguided baselines. Overall, our results show that small teacher models can reliably guide large generative systems, enabling stronger geometric control and broadening the creative range of text-to-image synthesis.

文本生成图像几何控制感知引导扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。