arXiv:2506.07977cs.CV2025-06NeurIPS被引 76

为图像生成模型设计多维度精细评估框架,解决现有评测不足问题。

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

  • 构建涵盖对齐、文本渲染等五个维度的综合评测体系
  • 支持按需选择子集评估,灵活高效完成模型分析
  • 适合研究者定位生成模型短板,推动技术迭代

文本到图像(T2I)模型在生成与文本提示一致的高质量图像方面受到广泛关注。然而,快速发展的T2I模型暴露出早期评测体系的局限性,尤其在推理能力、文本呈现和风格化等方面的评估不足。近期顶尖模型凭借强大的知识建模能力,在需要强推理的任务上表现优异,但现有评估系统尚未充分覆盖这一前沿。为此,我们提出OneIG-Bench,一个精心设计的综合性基准框架,用于对T2I模型在多个维度进行细粒度评估,包括提示-图像对齐、文本渲染精度、推理生成内容、风格化及多样性。通过结构化评估,该框架可深入分析模型性能,帮助研究人员和从业者定位图像生成全流程中的优势与瓶颈。具体而言,OneIG-Bench支持用户聚焦特定评估子集,仅针对选定维度的提示生成图像并完成相应评估,实现灵活高效。代码库与数据集已公开,便于复现研究与跨模型比较。

原文摘要 · Abstract (English)

Text-to-image (T2I) models have garnered significant attention for generating high-quality images aligned with text prompts. However, rapid T2I model advancements reveal limitations in early benchmarks, lacking comprehensive evaluations, for example, the evaluation on reasoning, text rendering and style. Notably, recent state-of-the-art models, with their rich knowledge modeling capabilities, show promising results on the image generation problems requiring strong reasoning ability, yet existing evaluation systems have not adequately addressed this frontier. To systematically address these gaps, we introduce OneIG-Bench, a meticulously designed comprehensive benchmark framework for fine-grained evaluation of T2I models across multiple dimensions, including prompt-image alignment, text rendering precision, reasoning-generated content, stylization, and diversity. By structuring the evaluation, this benchmark enables in-depth analysis of model performance, helping researchers and practitioners pinpoint strengths and bottlenecks in the full pipeline of image generation. Specifically, OneIG-Bench enables flexible evaluation by allowing users to focus on a particular evaluation subset. Instead of generating images for the entire set of prompts, users can generate images only for the prompts associated with the selected dimension and complete the corresponding evaluation accordingly. Our codebase and dataset are now publicly available to facilitate reproducible evaluation studies and cross-model comparisons within the T2I research community.

图像生成评估基准多维度评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。