通过迭代优化提示词,更公平地评估文生图模型真实能力
ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization
- 用多模态反馈循环优化提示词,分离提示与生成能力
- 优化后模型在组合生成上表现显著提升,部分能力被此前低估
- 优化提示可跨模型通用,揭示共性提示偏好
当前文生图基准测试使用固定提示词,因提示敏感性可能导致模型生成能力被低估,并产生偏向某些模型的偏差。本文提出ConceptMix++框架,通过迭代提示优化,将提示表达与视觉生成能力解耦。基于ConceptMix,该方法引入多模态优化流程,利用视觉语言模型反馈系统性地改进提示词。在多个扩散模型上的实验表明,优化后的提示显著提升组合生成性能,揭示了此前未被发现的模型能力,实现了更公平的文生图模型对比。分析显示,空间关系和形状等视觉概念在优化后收益更大,说明现有基准系统性低估了模型在这些方面的表现。此外,优化提示具有强跨模型迁移性,表明不同模型对有效提示词存在共性偏好。结果表明,僵化的基准测试可能严重低估真实模型能力,而本框架提供了更准确的评估方式与未来发展方向。
原文摘要 · Abstract (English)
Current text-to-image (T2I) benchmarks evaluate models on rigid prompts, potentially underestimating true generative capabilities due to prompt sensitivity and creating biases that favor certain models while disadvantaging others. We introduce ConceptMix++, a framework that disentangles prompt phrasing from visual generation capabilities by applying iterative prompt optimization. Building on ConceptMix, our approach incorporates a multimodal optimization pipeline that leverages vision-language model feedback to refine prompts systematically. Through extensive experiments across multiple diffusion models, we show that optimized prompts significantly improve compositional generation performance, revealing previously hidden model capabilities and enabling fairer comparisons across T2I models. Our analysis reveals that certain visual concepts -- such as spatial relationships and shapes -- benefit more from optimization than others, suggesting that existing benchmarks systematically underestimate model performance in these categories. Additionally, we find strong cross-model transferability of optimized prompts, indicating shared preferences for effective prompt phrasing across models. These findings demonstrate that rigid benchmarking approaches may significantly underrepresent true model capabilities, while our framework provides more accurate assessment and insights for future development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。