arXiv:2601.06600cs.CL2026-01ACL被引 1

测试中文短视频谣言中认知偏差对多模态大模型的影响

Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation

  • 构建200个视频的标注数据集,细粒度标注三类误导模式
  • Gemini-2.5-Pro在多模态下表现最佳,信念得分71.5/100
  • 模型易受权威账号等社会线索影响,存在认知偏差

短视频平台已成为虚假信息传播的主要渠道,其中欺骗性言论常利用视觉实验与社会线索。尽管多模态大语言模型(MLLMs)展现出强大的推理能力,其在认知偏差交织的虚假信息面前的鲁棒性仍待探索。本文提出一个综合性评估框架,基于200个短视频的高质量人工标注数据集,覆盖四个健康领域。该数据集对三种误导模式——实验错误、逻辑谬误和虚构主张——进行细粒度标注,并由国家标准和学术文献验证。我们在五种模态设置下评估八款前沿MLLMs。实验结果表明,Gemini-2.5-Pro在多模态设置下表现最优,信念得分为71.5/100,而o3最差,仅为35.2。此外,我们研究了引发错误信念的社会线索,发现模型易受权威频道标识等偏见影响。

原文摘要 · Abstract (English)

Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues. While Multimodal Large Language Models (MLLMs) have demonstrated impressive reasoning capabilities, their robustness against misinformation entangled with cognitive biases remains under-explored. In this paper, we introduce a comprehensive evaluation framework using a high-quality, manually annotated dataset of 200 short videos spanning four health domains. This dataset provides fine-grained annotations for three deceptive patterns-experimental errors, logical fallacies, and fabricated claims-each verified by evidence such as national standards and academic literature. We evaluate eight frontier MLLMs across five modality settings. Experimental results demonstrate that Gemini-2.5-Pro achieves the highest performance in the multimodal setting with a belief score of 71.5/100, while o3 performs the worst at 35.2. Furthermore, we investigate social cues that induce false beliefs in videos and find that models are susceptible to biases like authoritative channel IDs.

多模态模型认知偏差虚假信息短视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。