arXiv:2512.02794cs.CV2025-12

让AI生成图像时真实还原物体物理属性,如重量、弹性等。

PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation

  • 引入两种新损失函数,让模型学习物理概念并分离独立属性。
  • 在多个数据集上显著提升物理属性生成的准确率和视觉质量。
  • 适合需要真实物理模拟的领域,如工业设计、虚拟现实。

基于扩散模型的文本到图像定制方法在理解风格和形状等具体概念方面取得了显著进展,但对真实且复杂的物理概念定制仍关注不足。当前方法的核心缺陷在于训练过程中未显式引入物理知识。即使输入提示中包含与物理相关的词汇,实验表明这些方法仍无法在生成结果中准确体现相应物理属性。本文提出PhyCustom,一种包含两种新型正则化损失的微调框架,旨在激活扩散模型进行物理定制。具体而言,等距损失促使模型学习物理概念,解耦损失则消除独立概念间的混合学习。在多样化的数据集上开展实验,基准结果表明,PhyCustom在物理定制的定量和定性指标上均优于现有最先进方法。

原文摘要 · Abstract (English)

Recent diffusion-based text-to-image customization methods have achieved significant success in understanding concrete concepts to control generation processes, such as styles and shapes. However, few efforts dive into the realistic yet challenging customization of physical concepts. The core limitation of current methods arises from the absence of explicitly introducing physical knowledge during training. Even when physics-related words appear in the input text prompts, our experiments consistently demonstrate that these methods fail to accurately reflect the corresponding physical properties in the generated results. In this paper, we propose PhyCustom, a fine-tuning framework comprising two novel regularization losses to activate diffusion model to perform physical customization. Specifically, the proposed isometric loss aims at activating diffusion models to learn physical concepts while decouple loss helps to eliminate the mixture learning of independent concepts. Experiments are conducted on a diverse dataset and our benchmark results demonstrate that PhyCustom outperforms previous state-of-the-art and popular methods in terms of physical customization quantitatively and qualitatively.

物理生成扩散模型定制化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。