arXiv:2505.04650cs.GRcs.AI2025-05被引 5

构建文本生成图像的统一评测框架,提升提示词与模型选择效果

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models

  • 用结构化元数据增强提示词,提升生成图像质量
  • 多指标评估显示元数据显著提高真实感与语义准确性
  • 可指导特定任务下模型和提示词的优选,适合应用开发人员

本文提出一个开源的文本到图像生成模型统一评测框架,重点研究元数据增强提示词的影响。基于DeepFashion-MultiModal数据集,采用加权评分、基于CLIP的相似性、LPIPS、FID及检索类指标进行量化评估,并结合定性分析。结果表明,结构化元数据显著提升多种文本到图像架构下的视觉真实感、语义保真度与模型鲁棒性。虽非传统推荐系统,该框架可根据评估指标实现任务导向的模型与提示词推荐,助力实际应用中的选型决策。

原文摘要 · Abstract (English)

This work presents an open-source unified benchmarking and evaluation framework for text-to-image generation models, with a particular focus on the impact of metadata augmented prompts. Leveraging the DeepFashion-MultiModal dataset, we assess generated outputs through a comprehensive set of quantitative metrics, including Weighted Score, CLIP (Contrastive Language Image Pre-training)-based similarity, LPIPS (Learned Perceptual Image Patch Similarity), FID (Frechet Inception Distance), and retrieval-based measures, as well as qualitative analysis. Our results demonstrate that structured metadata enrichments greatly enhance visual realism, semantic fidelity, and model robustness across diverse text-to-image architectures. While not a traditional recommender system, our framework enables task-specific recommendations for model selection and prompt design based on evaluation metrics.

文本生成图像模型评测提示词优化元数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。