arXiv:2412.11767cs.CV2024-12CVPR被引 6

评测生成模型在真实设计任务中的表现,发现当前最佳模型仅达22.48分。

IDEA-Bench: How Far are Generative Models from Professional Designing?

  • 构建包含100个真实设计任务的综合评测基准IDEA-Bench
  • 顶尖模型在复杂任务中平均得分仅22.48,通用模型更低至6.81
  • 提供自动评估工具与在线排行榜,助力模型优化

真实世界的设计任务(如绘本创作、影视分镜、修图、视觉特效、字体迁移等)高度多样且复杂,需从指令、描述和参考图中深层解析多种元素。现有生成模型虽能根据提示生成高质量图像,但在涉及多输入输出、形式多样的专业设计场景中仍存在显著局限,即使使用ControlNets、LoRAs等增强方法亦然。为此,本文提出IDEA-Bench,一个涵盖100个真实设计任务的综合性基准,包括渲染、视觉特效、分镜、绘本、字体、风格化及身份保持生成等,共275个测试用例,全面评估模型的通用生成能力。结果显示,最优模型在该基准上仅得22.48分,通用模型最高仅6.81分。研究还分析了性能瓶颈,并提出改进方向。此外,提供18个代表性任务的多模态大模型自动评估方案,支持快速模型开发与对比。数据集、评估工具包与在线排行榜已开源:https://github.com/ali-vilab/IDEA-Bench,旨在推动生成模型向更通用、实用的智能设计系统演进。

原文摘要 · Abstract (English)

Real-world design tasks - such as picture book creation, film storyboard development using character sets, photo retouching, visual effects, and font transfer - are highly diverse and complex, requiring deep interpretation and extraction of various elements from instructions, descriptions, and reference images. The resulting images often implicitly capture key features from references or user inputs, making it challenging to develop models that can effectively address such varied tasks. While existing visual generative models can produce high-quality images based on prompts, they face significant limitations in professional design scenarios that involve varied forms and multiple inputs and outputs, even when enhanced with adapters like ControlNets and LoRAs. To address this, we introduce IDEA-Bench, a comprehensive benchmark encompassing 100 real-world design tasks, including rendering, visual effects, storyboarding, picture books, fonts, style-based, and identity-preserving generation, with 275 test cases to thoroughly evaluate a model's general-purpose generation capabilities. Notably, even the best-performing model only achieves 22.48 on IDEA-Bench, while the best general-purpose model only achieves 6.81. We provide a detailed analysis of these results, highlighting the inherent challenges and providing actionable directions for improvement. Additionally, we provide a subset of 18 representative tasks equipped with multimodal large language model (MLLM)-based auto-evaluation techniques to facilitate rapid model development and comparison. We releases the benchmark data, evaluation toolkits, and an online leaderboard at https://github.com/ali-vilab/IDEA-Bench, aiming to drive the advancement of generative models toward more versatile and applicable intelligent design systems.

生成模型设计评测多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。