arXiv:2510.14376cs.CV2025-10AAAI被引 2

通过方向性分离文本嵌入,提升多物体图像生成准确性。

DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation

  • 在CLIP文本嵌入中引入方向性分离机制
  • 多物体生成成功率显著提升,对象混淆减少
  • 适合需要精准多物体现的生成任务

文本到图像(T2I)生成模型虽在高质量图像生成上取得进展,但在多物体提示下仍易出现对象遗漏或混淆。通过系统研究,我们识别出四种典型失败场景:相似形状、相似纹理、背景偏差差异和多物体共现。基于对CLIP嵌入的两项关键观察,提出DOS(方向性对象分离)方法,通过修改三类CLIP文本嵌入来改进输入。实验显示,DOS持续提升多物体生成成功率达26.24%–43.04%(人类评估),显著优于四种对比方法,在四个基准测试中均表现更优。结果表明DOS是改善多物体图像生成的有效实用方案。

原文摘要 · Abstract (English)

Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation.

图像生成多物体CLIP文本嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。