arXiv:2503.16406cs.GRcs.CV2025-03CVPR被引 9

让AI更懂人与物互动细节,生成图像更准确。

VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness

  • 解耦互动词与常见词,增强对动作语义的理解
  • 在HICO-DET数据集上显著提升互动准确性
  • 无需额外条件,直接提升生成图像的语义对齐

当前大规模文本到图像扩散模型虽能生成逼真图像,但在刻画人与物体间的交互时往往表现不佳,因其难以区分不同交互词汇。本文提出VerbDiff,一种新型文本到图像生成模型,通过将交互词从基于频率的锚定词中解耦,并利用生成图像中的局部交互区域,增强模型对特定词汇语义的理解,从而减少交互词与物体间的偏差。该方法无需额外条件,即可使模型准确理解人与物之间的预期交互,生成高质量且语义对齐的图像。在HICO-DET数据集上的大量实验表明,该方法优于现有方法。

原文摘要 · Abstract (English)

Recent large-scale text-to-image diffusion models generate photorealistic images but often struggle to accurately depict interactions between humans and objects due to their limited ability to differentiate various interaction words. In this work, we propose VerbDiff to address the challenge of capturing nuanced interactions within text-to-image diffusion models. VerbDiff is a novel text-to-image generation model that weakens the bias between interaction words and objects, enhancing the understanding of interactions. Specifically, we disentangle various interaction words from frequency-based anchor words and leverage localized interaction regions from generated images to help the model better capture semantics in distinctive words without extra conditions. Our approach enables the model to accurately understand the intended interaction between humans and objects, producing high-quality images with accurate interactions aligned with specified verbs. Extensive experiments on the HICO-DET dataset demonstrate the effectiveness of our method compared to previous approaches.

文本生成图像生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。