arXiv:2509.11967cs.LGcs.CL2025-09被引 1

测试大模型对不同立场论据的接受度,发现权威来源易影响其立场。

MillStone: How Open-Minded Are LLMs?

  • 构建首个系统性评测模型立场受外部论据影响的基准MillStone。
  • 多数模型在争议问题上表现开放,但权威来源可轻易改变其立场。
  • 适合关注AI信息可信度与风险的研究者和开发者。

配备网络搜索、信息检索等代理能力的大语言模型正逐步取代传统搜索引擎。随着用户越来越多地依赖大模型获取各类话题的信息,包括有争议和存在分歧的问题,理解模型输出中的立场和观点如何受其信息来源文档的影响变得至关重要。本文提出MillStone,这是首个旨在系统衡量外部论据对大模型在争议性议题上立场影响的基准。我们对九个主流大模型应用MillStone,测量它们对相反立场论据的开放程度、模型间一致性、最具说服力的论据及其差异。总体发现,大模型在大多数议题上表现出开放性。然而,权威信息源能轻易改变模型立场,凸显了信息源选择的重要性,以及基于大模型的信息检索与搜索系统可能被操纵的风险。

原文摘要 · Abstract (English)

Large language models equipped with Web search, information retrieval tools, and other agentic capabilities are beginning to supplant traditional search engines. As users start to rely on LLMs for information on many topics, including controversial and debatable issues, it is important to understand how the stances and opinions expressed in LLM outputs are influenced by the documents they use as their information sources. In this paper, we present MillStone, the first benchmark that aims to systematically measure the effect of external arguments on the stances that LLMs take on controversial issues (not all of them political). We apply MillStone to nine leading LLMs and measure how ``open-minded'' they are to arguments supporting opposite sides of these issues, whether different LLMs agree with each other, which arguments LLMs find most persuasive, and whether these arguments are the same for different LLMs. In general, we find that LLMs are open-minded on most issues. An authoritative source of information can easily sway an LLM's stance, highlighting the importance of source selection and the risk that LLM-based information retrieval and search systems can be manipulated.

大模型评估立场分析信息可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。