arXiv:2603.25732cs.CV2026-03被引 5

构建商业视觉内容生成评估基准,填补真实设计需求空白

BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation

  • 构建覆盖5类文档的20项任务评估体系
  • 测试26个模型在文本渲染与布局控制上的表现
  • 适合设计自动化、AI生成工具研发者使用

近年来图像生成模型的应用已从美学图像扩展至实际视觉内容创作。然而,现有基准主要聚焦自然图像合成,未能系统评估模型在真实商业设计任务中所需的结构化与多约束要求。本文提出BizGenEval,一个面向商业视觉内容生成的系统性评估基准。该基准涵盖幻灯片、图表、网页、海报和科学图示五类典型文档,评估文本渲染、版式控制、属性绑定与基于知识的推理四大核心能力,形成20项多样化评估任务。基准包含400个精心设计的提示和8000条人工验证的检查清单问题,用于严格检验生成图像是否满足复杂的视觉与语义约束。我们在26个主流图像生成系统(包括领先商用API与开源模型)上进行大规模评测,结果揭示当前生成模型与专业视觉内容创作需求之间存在显著能力差距。我们希望BizGenEval能成为真实商业视觉内容生成的标准评估基准。

原文摘要 · Abstract (English)

Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly focus on natural image synthesis and fail to systematically evaluate models under the structured and multi-constraint requirements of real-world commercial design tasks. In this work, we introduce BizGenEval, a systematic benchmark for commercial visual content generation. The benchmark spans five representative document types: slides, charts, webpages, posters, and scientific figures, and evaluates four key capability dimensions: text rendering, layout control, attribute binding, and knowledge-based reasoning, forming 20 diverse evaluation tasks. BizGenEval contains 400 carefully curated prompts and 8000 human-verified checklist questions to rigorously assess whether generated images satisfy complex visual and semantic constraints. We conduct large-scale benchmarking on 26 popular image generation systems, including state-of-the-art commercial APIs and leading open-source models. The results reveal substantial capability gaps between current generative models and the requirements of professional visual content creation. We hope BizGenEval serves as a standardized benchmark for real-world commercial visual content generation.

视觉生成评估基准商业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。