arXiv:2608.09666cs.AI2026-08被引 2

用类人策略高效评估图像生成模型,10%时间达成传统方法效果。

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

论文配图:Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
图 1 · 摘自论文原文
  • 模仿人类快速判断,多轮迭代生成针对性样本与评估。
  • 仅需传统10%时间完成评估,结果与基准相当。
  • 支持自定义需求,适合研究者和开发者快速测试模型。

视觉生成模型虽已实现高质量图像与视频生成,但评估常需采样数百至数千个样本,计算成本高昂。现有方法依赖固定流程,难以适配用户需求,且结果缺乏解释。受人类通过少量样本快速形成模型能力判断的启发,我们提出评估代理框架(Evaluation Agent),采用类人策略实现高效、动态、多轮评估,提供可定制的详细分析。给定自然语言评估请求后,代理将任务分解为子维度,生成针对性提示,从目标模型采样图像或视频,调用合适评估工具,并根据观察证据迭代更新计划,覆盖预设基准维度与开放用户关切。该框架具备高效性、可提示性、可解释性及跨模型与工具的可扩展性。实验表明,评估时间降低至传统方法的10%,同时保持相似性能。我们进一步构建了开放评估代理(Open-EA),基于EA-CoT-10K——一个由多轮评估回放生成的历史条件化指令微调语料库,利用Qwen2.5-3B-Instruct训练得到30亿参数的本地规划核心(EA-3B),在保留原代理结构化推理、工具调用与摘要协议的同时,减少对专有后端的依赖。实验证明,基于API的代理在主流文本到图像/视频基准上表现良好,而Open-EA在四个域内与三个域外的T2V生成器家族中均有效,显示所学策略具有部分跨家族迁移能力。

原文摘要 · Abstract (English)

Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive. Existing evaluation methods also rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. Mimicking how humans quickly form impressions of a model's capabilities from only a few samples, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses. Given a natural-language evaluation request, the agent decomposes it into sub-aspects, generates targeted prompts, samples images or videos from the evaluated model, invokes suitable evaluation tools, and iteratively updates its plan from the observed evidence, covering both predefined benchmark dimensions and open-ended user concerns. The framework is thus efficient, promptable, explainable, and scalable across models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. We further introduce Open Evaluation Agent (Open-EA) by constructing EA-CoT-10K, a corpus of history-conditioned step-level instruction-tuning records derived from multi-round evaluation rollouts, and training EA-3B from Qwen2.5-3B-Instruct as a local planning backbone that preserves the structured reasoning, tool invocation, and summary protocol of the API-based agent while reducing dependence on proprietary backbones. Experiments validate the API-based agent on established T2I/T2V benchmarks and open-ended queries, and evaluate Open-EA on four in-domain and three out-of-domain T2V generator families, showing partial cross-family transfer of the learned policy.

模型评估生成模型智能代理高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。