用诊断式提示提升大模型对亚马逊搜索页视觉复杂度的判断能力
Exploring Diagnostic Prompting Approach for Multimodal LLM-based Visual Complexity Assessment: A Case Study of Amazon Search Result Pages
- 采用诊断式提示分析视觉元素权重,引导模型聚焦设计细节
- F1得分提升858%,但整体准确率仍较低(κ=0.071)
- 适合关注人机评估对齐的大模型应用研究者
本研究探究诊断式提示能否提升多模态大模型(MLLM)在亚马逊搜索结果页(SRP)视觉复杂度评估中的可靠性。对比标准格式原则提示与诊断式提示,在200个亚马逊SRP页面和人工专家标注数据上进行评估。诊断式提示使预测人类复杂度判断的F1得分从0.031提升至0.297(相对提升858%),但绝对性能仍不理想(Cohen's κ = 0.071)。决策树分析显示,模型更重视视觉设计元素(如徽章杂乱度占比38.6%),而人类更关注内容相似性,说明两者推理模式存在部分重叠。故障案例分析表明,模型在产品相似性和颜色强度判断上仍存持续挑战。研究认为诊断式提示是实现人机对齐评估的有前景起点,但需更大规模真实标注数据支持,以克服人机一致性差的问题。
原文摘要 · Abstract (English)
This study investigates whether diagnostic prompting can improve Multimodal Large Language Model (MLLM) reliability for visual complexity assessment of Amazon Search Results Pages (SRP). We compare diagnostic prompting with standard gestalt principles-based prompting using 200 Amazon SRP pages and human expert annotations. Diagnostic prompting showed notable improvements in predicting human complexity judgments, with F1-score increasing from 0.031 to 0.297 (+858\% relative improvement), though absolute performance remains modest (Cohen's $κ$ = 0.071). The decision tree revealed that models prioritize visual design elements (badge clutter: 38.6\% importance) while humans emphasize content similarity, suggesting partial alignment in reasoning patterns. Failure case analysis reveals persistent challenges in MLLM visual perception, particularly for product similarity and color intensity assessment. Our findings indicate that diagnostic prompting represents a promising initial step toward human-aligned MLLM-based evaluation, though failure cases with consistent human-MLLM disagreement require continued research and refinement in prompting approaches with larger ground truth datasets for reliable practical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。