arXiv:2607.19355cs.AIcs.CL2026-07被引 1

测试大模型对信息来源和真实性的判断能力,发现其表现接近随机。

Information Discernment in Large Language Models

论文配图:Information Discernment in Large Language Models
图 1 · 摘自论文原文
  • 提出可衡量信息辨识能力的框架L2D,基于三个规范性公理
  • 13个模型在近67万次测试中,对可靠来源和真相更新均表现不佳
  • 新模型虽提升真相判断力,但来源辨识仍是短板,适合对齐研究者

大型语言模型日益依赖外部知识源(如互联网)。它们是否能合理权衡信息——即更信任可靠来源(来源辨识),并在事实更接近真相时做出更大调整(真相辨识)?本文将此问题形式化为“信息辨识”,并提出学习辨识(Learn2Discern, L2D)实验框架与基准,基于三个规范性公理设计可解释度量。为验证外部有效性,一项预注册、配额匹配的用户研究(n=299)表明真实用户支持所有公理,并报告违反这些原则会降低信任度和使用意愿。在13个模型与近67万次试验中,发现两者维度均存在系统性失败:模型在来源与真相辨识上表现接近随机,对来源流行度的依赖是可靠性权重的两倍,且无论信息改善或恶化其立场,更新幅度基本相同。模型仅在先验最准确的数据集上最有效地整合外部知识。较新、更大的模型提升了真相辨识能力,但未改善来源辨识,说明模型复杂度无法解决此盲区。本文还识别出简单推理时干预方法可同时提升两类辨识能力。数据集与调查问卷已公开,作为核心对齐属性的测试平台,随着大模型取代传统搜索,该属性重要性持续上升。

原文摘要 · Abstract (English)

LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (truth discernment)? We formalize this as information discernment and introduce Learn2Discern (L2D), an experimental framework and benchmark grounded in three normative axioms with interpretable metrics. To establish external validity, a pre-registered, quota-matched user study (n=299) confirms that real LLM users endorse all three axioms and report that violations reduce their trust and usage intent. Across 13 models and nearly 670K trials, we find consistent failures across both dimensions: models perform near chance on source and truth discernment, rely on source popularity twice as much as source reliability, and update roughly equally whether a claim improves or worsens their position relative to the ground truth. Models integrate external knowledge most effectively on datasets where their priors are already the most accurate. Newer and larger models improve truth discernment but not source discernment, a blind spot that model complexity does not address. We identify simple inference-time interventions that improve both forms of discernment. We release our dataset and survey as a testbed for a core alignment property that scales in importance as LLMs replace traditional search.

大模型对齐信息可信度推理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。