arXiv:2508.02074cs.CLcs.LG2025-08

测试大模型识别虚假信息的能力,发现主流模型仍易被误导。

The SMeL Test: A simple benchmark for media literacy in language models

  • 设计极简基准测试SMeL,评估模型过滤不可信信息的能力。
  • 最优秀模型仍有70%概率产生幻觉,且大模型不必然更优。
  • 适合关注AI可信度、内容安全的研究者和开发者参考。

互联网充斥着未署名、故意误导或不可信的内容。尽管大型语言模型常被用于自主网页浏览,但它们是否掌握了人类研究人员用于应对这种嘈杂环境的简单启发式方法尚不清楚。本文提出合成媒体素养测试(SMeL Test),一个最小化基准,用于测试语言模型在上下文中主动过滤不可信信息的能力。我们对多种常用指令微调的LLM进行了评估,包括推理型模型,发现没有模型能持续成功;尽管推理能力与更高得分相关,但即使表现最好的API模型,仍有高达70%的概率产生幻觉。值得注意的是,更大、更强大的模型并不一定优于较小模型。我们希望本工作能加深对该类幻觉的理解,并指导新方法的发展以应对该问题。

原文摘要 · Abstract (English)

The internet is rife with unattributed, deliberately misleading, or otherwise untrustworthy content. Though large language models (LLMs) are often tasked with autonomous web browsing, the extent to which they have learned the simple heuristics human researchers use to navigate this noisy environment is not currently known. In this paper, we introduce the Synthetic Media Literacy Test (SMeL Test), a minimal benchmark that tests the ability of language models to actively filter out untrustworthy information in context. We benchmark a variety of commonly used instruction-tuned LLMs, including reasoning models, and find that no model consistently succeeds; while reasoning in particular is associated with higher scores, even the best API model we test hallucinates up to 70% of the time. Remarkably, larger and more capable models do not necessarily outperform their smaller counterparts. We hope our work sheds more light on this important form of hallucination and guides the development of new methods to combat it.

媒体素养幻觉检测LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。