直接问模型是否传播,不如问可信度更能预测真实传播行为。
When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation
- 用模型评估虚假内容可信度比直接问传播意愿更准
- 290篇假文章测试显示,可信度评分预测传播率更优
- 提问方式比提问技巧更重要,间接评估可能更有效
大型语言模型(LLM)可大规模生成欺骗性内容,亟需可扩展的虚假信息风险评估方法,以判断读者是否认为内容可信并愿意分享。一种自然做法是直接向模型提问,将其返回的评分视为人类评分的预测值。该做法隐含假设:询问目标反应能获得最佳预测得分。我们基于317名参与者对290篇欺骗性文章的匹配可信度与分享意愿评分,以及8个LLM评估者的结果进行检验。结果显示:该假设在可信度上成立,但在分享意愿上失败。所有评估者中,可信度评分对人类分享行为的预测能力均不弱于甚至优于分享评分;而在预测新场景下的行为时,分享评分未带来显著提升。这一现象在问题顺序不同或分开展示时依然存在。结果表明,直接询问目标反应未必最有效,通过相关判断的间接路径可能更具预测力。选择问什么,可能比如何问同样关键。
原文摘要 · Abstract (English)
LLMs make it increasingly easy to generate deceptive content at scale, creating a need for scalable misinformation risk evaluation based on whether readers find such content credible and are willing to share it. A natural approach is to ask an LLM these questions directly and treat the returned scores as predictions of the corresponding human ratings. Implicit in this practice is the assumption that asking about a reader response produces the score that best predicts it. We test this assumption using matched credibility and willingness-to-share ratings for 290 deceptive articles from 317 participants and eight LLM evaluators. Unexpectedly, the assumption holds for credibility but fails for sharing. For every evaluator, credibility scores track human sharing at least as closely as sharing scores, while sharing scores offer no detectable benefit beyond credibility when predicting responses to unseen scenarios. This pattern persists when the questions are asked separately or in reversed order. Our results suggest that directly asking for the target response may not always yield the most effective score for predicting it. Comparing direct scores with indirect paths through related judgments may reveal a more effective predictive route. Deciding what to ask may be as important as refining how to ask it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。