首个将生成图像与真实商业设计付款挂钩的基准,评估模型在实际场景中的经济价值。
ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services

- 构建包含1070个真实付费设计任务的数据集,覆盖人像、产品等场景。
- 提出三维度评分体系,准确预测82%的人类付款决策。
- 适合关注生成模型商业化落地的研究者和设计师参考。
近期图像生成与编辑模型在学术基准上表现优异,但在真实商业项目中的效果尚不明确。本文提出【ServImage】,一个将模型输出与商业项目经济价值直接关联的基准。该基准包含:(i) 【ServImageBench】:包含1070个真实付费商业设计任务及2050份设计师交付成果,总价值超29.5万美元,涵盖人像、产品和数字内容,并配有3.3万张候选图与3.3万条人工标注;(ii) 【ServImageScore】:融合基础需求满足度、视觉执行质量与商业必要性三个维度的评分系统,反映人类付款决策的关键因素;(iii) 【ServImageModel】:基于人工标注数据训练的支付预测模型,在预测人类付款行为上达到82.00%准确率,输出校准后的支付概率。ServImage为评估图像生成模型的商业可行性提供了全面基础,是未来面向经济驱动视觉系统的可扩展研究资源。
原文摘要 · Abstract (English)
Recent image generation and editing models demonstrate robust adherence to instructions and high visual quality on academic benchmarks. However, their performance on paid, real-world design projects remains uncertain. We introduce \textbf{ServImage}, a benchmark that explicitly correlates model outputs with economic value in commercial design projects. ServImage consists of (i) \textbf{\textit{ServImageBench}}: a dataset of 1.07k paid commercial design tasks and 2.05k designer deliverables totaling over \$295k, covering portrait, product, and digital content, along with 33k candidate images and 33k human annotations. (ii) \textbf{\textit{ServImageScore}}: an integrated scoring system that combines three quality dimensions: baseline requirements fulfilment, visual execution quality, and commercial necessity satisfaction. These three dimensions are designed to characterize the factors that drive human payment decisions and indicate whether an image is commercially acceptable. (iii) \textbf{\textit{ServImageModel}}: under this scoring system, we propose a payment prediction model trained on the human-annotated candidate images, achieving 82.00\% accuracy in predicting human payment decisions and producing calibrated payment probabilities. ServImage provides a comprehensive foundation for assessing the commercial viability of image generation models and offers a scalable resource for future research on economically grounded vision systems \href{https://github.com/FengxianJi/ServImage}{Github.}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。