用大模型发现视觉内容可信度关键特征,提升预测与解释能力
Large Language Model-Informed Feature Discovery Improves Prediction and Interpretation of Credibility Perceptions of Visual Content
- 通过提示词引导大模型识别可解释的视觉特征
- 在8个主题4191条数据上使预测准确率提升13%(R²)
- 适合研究虚假信息、人机判断或大模型辅助社会科学研究者
在视觉主导的社交媒体环境中,预测视觉内容的可信度并理解人类判断依据对遏制虚假信息至关重要。然而,视觉特征的多样性和丰富性带来了挑战。本文提出一种基于大语言模型(LLM)的特征发现框架,利用GPT-4o等多模态大模型评估内容可信度并解释推理过程。通过针对性提示词提取并量化可解释特征,集成至机器学习模型以提升可信度预测性能。在涵盖科学、健康和政治共8个主题的4,191条社交媒体图文上进行测试,使用5,355名众包工作者提供的可信度评分。该方法相较零样本GPT预测在R²上提升13%,揭示了信息具体性和图像格式等关键特征。研究讨论了其在虚假信息治理、视觉可信度分析及大模型在社会科学中应用的启示。
原文摘要 · Abstract (English)
In today's visually dominated social media landscape, predicting the perceived credibility of visual content and understanding what drives human judgment are crucial for countering misinformation. However, these tasks are challenging due to the diversity and richness of visual features. We introduce a Large Language Model (LLM)-informed feature discovery framework that leverages multimodal LLMs, such as GPT-4o, to evaluate content credibility and explain its reasoning. We extract and quantify interpretable features using targeted prompts and integrate them into machine learning models to improve credibility predictions. We tested this approach on 4,191 visual social media posts across eight topics in science, health, and politics, using credibility ratings from 5,355 crowdsourced workers. Our method outperformed zero-shot GPT-based predictions by 13 percent in R2, and revealed key features like information concreteness and image format. We discuss the implications for misinformation mitigation, visual credibility, and the role of LLMs in social science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。