用少量样本快速评估图像生成模型,像人一样灵活又可解释。
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
- 模仿人类看样例的直觉,每轮只用少量样本动态评估。
- 评估时间仅需传统方法的10%,结果仍具可比性。
- 支持定制化提问,适合研究者和开发者快速测试模型。
视觉生成模型近年来在高质量图像与视频生成方面取得进展,但评估常需采样数百甚至上千个样本,计算成本高,尤其对采样慢的扩散模型更不友好。现有评估方法依赖固定流程,难以适应用户需求,且仅输出数值结果而缺乏解释。相比之下,人类只需观察少数样本即可形成判断。为此,我们提出Evaluation Agent框架,采用类人策略,通过少量样本实现高效、动态、多轮评估,并提供针对用户需求的详细分析。该框架具备四大优势:效率高、可按提示定制评估、超越单一分数的可解释性、可扩展至多种模型与工具。实验表明,Evaluation Agent将评估时间压缩至传统方法的10%的同时保持结果可比性。该框架已开源,旨在推动视觉生成模型及其高效评估的研究。
原文摘要 · Abstract (English)
Recent advancements in visual generative models have enabled high-quality image and video generation, opening diverse applications. However, evaluating these models often demands sampling hundreds or thousands of images or videos, making the process computationally expensive, especially for diffusion-based models with inherently slow sampling. Moreover, existing evaluation methods rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. In contrast, humans can quickly form impressions of a model's capabilities by observing only a few samples. To mimic this, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations using only a few samples per round, while offering detailed, user-tailored analyses. It offers four key advantages: 1) efficiency, 2) promptable evaluation tailored to diverse user needs, 3) explainability beyond single numerical scores, and 4) scalability across various models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. The Evaluation Agent framework is fully open-sourced to advance research in visual generative models and their efficient evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。