探究视觉语言模型能否理解视觉说服力,发现其核心短板是难懂意图。
Do Vision-Language Models Understand Visual Persuasiveness?
- 构建高共识二元说服力判断数据集,提出三层次视觉说服因素分类。
- 模型对高说服力过度预测,低/中层特征区分能力弱,语义对齐最关键。
- 具体对象导向的解释能显著提升准确率,盲目提示反而降低效果。
视觉语言模型(VLMs)在多模态推理与理解方面取得显著进展,但其是否真正掌握视觉说服力——即视觉线索如何影响人类态度与决策——仍不明确。为此,我们构建了一个高共识的二元说服力判断数据集,并提出视觉说服因素(VPFs)分类体系,涵盖低层感知、中层构图和高层语义线索。我们还探索了认知引导与知识注入等说服相关推理策略。跨VLM的实证分析显示,模型存在回忆导向偏差,过度预测高说服力,且对低/中层特征的判别力弱;而信息与对象存在的高层语义对齐成为人类判断的最强预测因子。干预策略中,简单指令或无指导推理框架效果微弱甚至负面,而简洁、以对象为基础的推理理由显著提升精度与F1分数。结果表明,VLM的核心局限并非识别说服对象,而在于将其与传播意图关联。
原文摘要 · Abstract (English)
Recent advances in vision-language models (VLMs) have enabled impressive multi-modal reasoning and understanding. Yet, whether these models truly grasp visual persuasion-how visual cues shape human attitudes and decisions-remains unclear. To probe this question, we construct a high-consensus dataset for binary persuasiveness judgment and introduce the taxonomy of Visual Persuasive Factors (VPFs), encompassing low-level perceptual, mid-level compositional, and high-level semantic cues. We also explore cognitive steering and knowledge injection strategies for persuasion-relevant reasoning. Empirical analysis across VLMs reveals a recall-oriented bias-models over-predict high persuasiveness-and weak discriminative power for low/mid-level features. In contrast, high-level semantic alignment between message and object presence emerges as the strongest predictor of human judgment. Among intervention strategies, simple instruction or unguided reasoning scaffolds yield marginal or negative effects, whereas concise, object-grounded rationales significantly improve precision and F1 scores. These results indicate that VLMs core limitation lies not in recognizing persuasive objects but in linking them to communicative intent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。