构建海报理解与生成的综合性设计评测基准,推动模型更懂设计意图。
PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation
- 基于版式、字体层级和语义意图三维度标注7765个图文实例
- 发现现有模型在视觉层次与设计意图表达上存在明显短板
- 适合研究生成式AI与人本设计融合的学者与开发者
我们提出PosterIQ,一个以设计为导向的海报理解与生成评测基准,涵盖版式结构、文字层级和语义意图三方面的标注。数据集包含7,765个图像-注释实例和822个生成提示,覆盖真实、专业及合成场景。为连接视觉设计认知与生成建模,我们定义了布局解析、图文对应、字体可读性与感知、设计质量评估,以及含隐喻的可控、结构感知生成等任务。评估主流多模态大模型与扩散生成器发现:在视觉层次、字体语义、显著性控制和意图传达方面仍存在显著差距;商业模型在高层推理上表现较好,但对设计判断缺乏敏感度,仅作机械评分;生成器虽能良好渲染文字,却难以实现结构化合成。深入分析表明,PosterIQ既是量化评测工具,也是诊断设计推理能力的工具,提供可复现的任务特异性指标。旨在推动模型创造力提升,并将以人为本的设计原则融入生成式视觉语言系统。
原文摘要 · Abstract (English)
We present PosterIQ, a design-driven benchmark for poster understanding and generation, annotated across composition structure, typographic hierarchy, and semantic intent. It includes 7,765 image-annotation instances and 822 generation prompts spanning real, professional, and synthetic cases. To bridge visual design cognition and generative modeling, we define tasks for layout parsing, text-image correspondence, typography/readability and font perception, design quality assessment, and controllable, composition-aware generation with metaphor. We evaluate state-of-the-art MLLMs and diffusion-based generators, finding persistent gaps in visual hierarchy, typographic semantics, saliency control, and intention communication; commercial models lead on high-level reasoning but act as insensitive automatic raters, while generators render text well yet struggle with composition-aware synthesis. Extensive analyses show PosterIQ is both a quantitative benchmark and a diagnostic tool for design reasoning, offering reproducible, task-specific metrics. We aim to catalyze models' creativity and integrate human-centred design principles into generative vision-language systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。