arXiv:2409.08251cs.CV2024-09中稿 · ACM MM 2024被引 6

动态调整提示词,让冻结的文生图模型更精准定位图像中的物体。

Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding

  • 在扩散模型UNet中设计动态提示更新模块,融合图像特征优化文本提示。
  • 在PNG基准上达到新最好性能,显著提升细粒度对齐能力。
  • 适合需要高精度图文定位的视觉语言任务研究者使用。

全景叙事定位(PNG)的核心目标是实现细粒度的图文对齐,即根据叙事描述对图像中的目标进行全景分割。以往判别式方法通过全景分割预训练或CLIP模型适配,仅能实现弱或粗粒度对齐。随着文生图扩散模型的发展,已有研究证明其可通过交叉注意力图和改进的通用分割性能实现细粒度对齐。然而,直接使用短语特征作为静态提示来应用冻结的扩散模型仍存在较大任务差距和不足的跨模态交互,导致性能不佳。为此,我们提出一种嵌入于扩散模型UNet中的可提取-可注入短语适配器(EIPA),动态利用图像特征更新短语提示,并将多模态线索回注,更充分地挖掘扩散模型的细粒度对齐能力。此外,还设计了多层级互聚合(MLMA)模块,以递归融合多层级图像与短语特征,实现分割优化。在PNG基准上的大量实验表明,该方法实现了新的最佳性能。

原文摘要 · Abstract (English)

Panoptic narrative grounding (PNG), whose core target is fine-grained image-text alignment, requires a panoptic segmentation of referred objects given a narrative caption. Previous discriminative methods achieve only weak or coarse-grained alignment by panoptic segmentation pretraining or CLIP model adaptation. Given the recent progress of text-to-image Diffusion models, several works have shown their capability to achieve fine-grained image-text alignment through cross-attention maps and improved general segmentation performance. However, the direct use of phrase features as static prompts to apply frozen Diffusion models to the PNG task still suffers from a large task gap and insufficient vision-language interaction, yielding inferior performance. Therefore, we propose an Extractive-Injective Phrase Adapter (EIPA) bypass within the Diffusion UNet to dynamically update phrase prompts with image features and inject the multimodal cues back, which leverages the fine-grained image-text alignment capability of Diffusion models more sufficiently. In addition, we also design a Multi-Level Mutual Aggregation (MLMA) module to reciprocally fuse multi-level image and phrase features for segmentation refinement. Extensive experiments on the PNG benchmark show that our method achieves new state-of-the-art performance.

图文对齐扩散模型视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。