arXiv:2508.19289cs.CVcs.AI2025-08

用设计提示增强无监督评估,让系统像专家一样判断幻灯片质量。

Seeing Like a Designer Without One: A Study on Unsupervised Slide Quality Assessment via Designer Cue Augmentation

  • 融合7项设计指标与CLIP-ViT特征,用异常检测评估幻灯片质量。
  • 与人工评分相关性达0.83,优于主流视觉语言模型1.79至3.23倍。
  • 适合需要实时、客观幻灯片反馈的教学或演讲优化场景。

我们提出一种无监督幻灯片质量评估流程,结合七项专家启发的视觉设计指标(留白、色彩丰富度、边缘密度、亮度对比、文本密度、色彩和谐、版式平衡)与CLIP-ViT嵌入,采用基于孤立森林的异常评分机制评估演示文稿。该方法在1.2万张专业讲座幻灯片上训练,并在六场学术演讲(共115张幻灯片)上评估,与人工视觉评分的皮尔逊相关系数最高达0.83,比主流视觉语言模型(ChatGPT o4-mini-high、ChatGPT o3、Claude Sonnet 4、Gemini 2.5 Pro)得分强1.79至3.23倍。结果表明,该方法在视觉评分上具有收敛效度,在演讲表现评分上具备区分效度,并初步匹配整体印象。研究证明,通过多模态嵌入增强底层设计线索,可精准逼近观众对幻灯片质量的感知,实现可扩展的实时客观反馈。

原文摘要 · Abstract (English)

We present an unsupervised slide-quality assessment pipeline that combines seven expert-inspired visual-design metrics (whitespace, colorfulness, edge density, brightness contrast, text density, color harmony, layout balance) with CLIP-ViT embeddings, using Isolation Forest-based anomaly scoring to evaluate presentation slides. Trained on 12k professional lecture slides and evaluated on six academic talks (115 slides), our method achieved Pearson correlations up to 0.83 with human visual-quality ratings-1.79x to 3.23x stronger than scores from leading vision-language models (ChatGPT o4-mini-high, ChatGPT o3, Claude Sonnet 4, Gemini 2.5 Pro). We demonstrate convergent validity with visual ratings, discriminant validity against speaker-delivery scores, and exploratory alignment with overall impressions. Our results show that augmenting low-level design cues with multimodal embeddings closely approximates audience perceptions of slide quality, enabling scalable, objective feedback in real time.

幻灯片评估无监督学习CLIP设计指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。