arXiv:2606.00592cs.CV2026-06

提出可解释的多尺度视觉设计评估框架,让AI像人一样理解设计原则。

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs

论文配图:Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs
图 1 · 摘自论文原文
  • 构建10万+样本的可控制设计缺陷数据集,分离每项设计原则
  • 发现主流模型对特定设计问题不敏感,仅大模型有全局感知
  • 融合量化评分、局部反馈与全局推理,实现可解释的改进建议

有效的视觉传达依赖于可读性、对比度、对齐、重叠和连贯性等多重设计原则的和谐统一。尽管人类设计师能整体把握这些原则,机器代理通常将其简化为单一评分,缺乏可解释性和诊断精度。为此,我们提出PRISM(原则感知、可解释、结构引导的设计修改基准),系统性地对Crello数据集中的专业版面进行可测量的设计原则扰动。该基准包含10万条扰动训练样本和1万条扰动验证设计,每项均隔离特定原则的违反情况,以支持多模态设计质量推理的受控分析。我们发现Qwen-2.5-VL和GPT-4o-mini对目标原则退化反应迟钝,而GPT-4o虽具全局意识但缺乏细粒度解耦。基于此,我们提出一种多尺度评估框架,整合轻量级评分器用于定量评估、指令微调的视觉语言模型提供局部反馈、提示工程方法支持全局推理。该框架可生成可解释的设计失败原因说明,并利用局部洞察实现针对性优化。PRISM与该框架共同奠定可解释设计认知型多模态推理系统的基础。

原文摘要 · Abstract (English)

Effective visual communication stems from the harmony of multiple design principles, such as readability, contrast, alignment, overlap, and coherence, which collectively govern clarity and intent of the communicator. While human designers reason holistically over these principles, machine agents typically condense them into a single heuristic score, offering limited interpretability and diagnostic precision. To address this gap, we introduce PRISM (PRinciple-aware, Interpretable, and Structure-guided Design Modifications), a benchmark that systematically perturbs professional layouts from the Crello dataset along measurable design principles. The benchmark comprises 100K perturbed training samples and 10K perturbed validation designs, each isolating a specific principle violation for controlled analysis of multimodal reasoning about design quality. We show that models like Qwen-2.5-VL and GPT-4o-mini are largely insensitive to targeted principle degradations, whereas GPT-4o exhibits global awareness without fine-grained disentanglement. Building on these insights, we propose a multi-scale evaluation framework that integrates lightweight scorers for quantitative assessment, instruction-tuned vision-language models for localised feedback, and prompt-based methods for global reasoning. Our framework provides interpretable explanations of design failures. Using these localised insights, we show targeted refinements that improve layout quality. Together, PRISM and our framework lay the foundation for interpretable design-literate multimodal reasoning systems.

视觉设计可解释AI多模态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。