arXiv:2510.15021cs.CV2025-10

用社交媒体真实使用数据构建新评测基准,更贴近模型实际应用。

Constantly Improving Image Models Need Constantly Improving Benchmarks

  • 从用户社交帖子中提取真实提示,构建动态评测集
  • 发现31,000条现有基准未覆盖的复杂任务
  • 结合社区反馈设计更贴合实际的质量评估指标

近期图像生成技术(如GPT-4o Image Gen)不断涌现新能力,重塑用户交互方式。但现有评测基准常滞后,难以捕捉新兴应用场景,导致社区感知与正式评估之间存在差距。为此,我们提出ECHO框架,通过分析社交媒体上展示新型提示和用户评价的真实内容,构建评测基准。以GPT-4o Image Gen为例,我们收集了超过31,000条来自社交帖子的提示。分析表明,ECHO(1)发现了现有基准缺失的创造性复杂任务,如跨语言重绘产品标签或生成指定金额的收据;(2)能更清晰区分顶尖模型与替代方案;(3)通过社区反馈指导设计质量度量指标,例如检测颜色、身份和结构的显著变化。项目官网:https://echo-bench.github.io。

原文摘要 · Abstract (English)

Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture these emerging use cases, leaving a gap between community perceptions of progress and formal evaluation. To address this, we present ECHO, a framework for constructing benchmarks directly from real-world evidence of model use: social media posts that showcase novel prompts and qualitative user judgments. Applying this framework to GPT-4o Image Gen, we construct a dataset of over 31,000 prompts curated from such posts. Our analysis shows that ECHO (1) discovers creative and complex tasks absent from existing benchmarks, such as re-rendering product labels across languages or generating receipts with specified totals, (2) more clearly distinguishes state-of-the-art models from alternatives, and (3) surfaces community feedback that we use to inform the design of metrics for model quality (e.g., measuring observed shifts in color, identity, and structure). Our website is at https://echo-bench.github.io.

图像生成评测基准用户数据真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。