arXiv:2504.07046cs.CVcs.CL2025-04ACL被引 11

用大模型自动评估条件图像生成,效果接近人工评判。

A Unified Agentic Framework for Evaluating Conditional Image Generation

  • 以多模态大模型为核心,构建可自主选择工具的评估框架。
  • 在7个任务上与人工评估相关性达0.4625,接近人工间相关性0.47。
  • 仅需2.3K训练轨迹即可在小模型上超越此前GPT-4o最优方法。

条件图像生成因其个性化内容生成能力受到广泛关注,但缺乏任务无关、可靠且可解释的评估指标。本文提出CIGEval,一种统一的智能体式评估框架,用于全面评估条件图像生成任务。该框架以大视觉语言模型(LMM)为核心,集成多功能工具箱并建立细粒度评估体系。我们还设计了评估轨迹用于微调,使小型LMM能自主选择工具并基于输出进行精细分析。在七个主流条件图像生成任务上的实验表明,采用GPT-4o版本的CIGEval与人工评估的相关性高达0.4625,接近人工标注者间相关性0.47。当使用仅2.3K训练轨迹的7B开源LMM实现时,CIGEval性能超越此前基于GPT-4o的最先进方法。对GPT-4o生成图像的案例研究显示,该框架能有效识别主体一致性及控制指令遵循中的细微问题,展现出达到人类水平可靠性自动化评估的巨大潜力。

原文摘要 · Abstract (English)

Conditional image generation has gained significant attention for its ability to personalize content. However, the field faces challenges in developing task-agnostic, reliable, and explainable evaluation metrics. This paper introduces CIGEval, a unified agentic framework for comprehensive evaluation of conditional image generation tasks. CIGEval utilizes large multimodal models (LMMs) as its core, integrating a multi-functional toolbox and establishing a fine-grained evaluation framework. Additionally, we synthesize evaluation trajectories for fine-tuning, empowering smaller LMMs to autonomously select appropriate tools and conduct nuanced analyses based on tool outputs. Experiments across seven prominent conditional image generation tasks demonstrate that CIGEval (GPT-4o version) achieves a high correlation of 0.4625 with human assessments, closely matching the inter-annotator correlation of 0.47. Moreover, when implemented with 7B open-source LMMs using only 2.3K training trajectories, CIGEval surpasses the previous GPT-4o-based state-of-the-art method. Case studies on GPT-4o image generation highlight CIGEval's capability in identifying subtle issues related to subject consistency and adherence to control guidance, indicating its great potential for automating evaluation of image generation tasks with human-level reliability.

图像评估大模型应用自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。