arXiv:2605.08965cs.CV2026-05

测试大模型能否靠谱地解释图像为何有说服力。

Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning

论文配图:Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
图 1 · 摘自论文原文
  • 用多样化的教师生成推理文本微调模型,提升说服力预测效果。
  • 发现仅靠提示推理反而可能降低性能,且推理与判断不一致。
  • 提出三维评估框架,适合研究可信推理的AI系统设计者。

尽管多模态大语言模型在多模态任务上表现优异,但预测图像是否具有说服力及其原因仍具挑战性。我们首先表明,提示模型在预测前进行推理并未持续提升性能,甚至会降低说服力预测准确率,说明盲目生成的推理不可靠。目前尚无有效方法训练模型推理视觉说服力,也缺乏评估其推理是否忠实的方法。为此,我们通过实证与理论分析发现,使用多样化教师生成的推理进行监督微调可显著提升视觉说服力预测能力。我们进一步提出一个三维忠实度评估框架,涵盖推理与决策一致性、推理与图像的关联性以及推理对决策的敏感性。实验表明,仅凭预测性能无法保证推理的忠实性,其中推理对决策的敏感性最符合人类对推理偏好。这些发现推动了面向忠实性的训练目标与可扩展的推理监督机制。代码与数据集将公开发布。

原文摘要 · Abstract (English)

Despite strong performance of Multimodal Large Language Models (MLLMs) on multimodal tasks, predicting whether and why an image is persuasive remains challenging. We first show that prompting MLLMs to reason before prediction does not consistently help, and can even reduce persuasiveness prediction performance, suggesting that naively generated rationales are unreliable signals for this task. Yet, no established methodology exists for training MLLMs to reason about visual persuasion or evaluating whether their rationales faithfully support their decisions. To address this gap, we show empirically and theoretically that diverse teacher-generated rationales, when used for supervised fine-tuning, improve visual persuasiveness prediction. We further introduce a three-dimensional faithfulness evaluation framework covering rationale-to-decision consistency, rationale-to-image groundedness, and rationale-to-decision sensitivity. Applying this framework shows that prediction performance alone does not guarantee faithful rationales, while rationale-to-decision sensitivity is most aligned with human rationale preferences. These findings motivate faithfulness-aware training objectives and scalable rationale supervision for visual persuasiveness evaluation. Our code and dataset will be made publicly available.

多模态推理可信度说服力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。