arXiv:2607.20487cs.AIcs.CL2026-07

发现大模型在政治问答中普遍存在向左偏移的虚构内容。

Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering

论文配图:Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering
图 1 · 摘自论文原文
  • 用真实新闻生成问题,检测模型回答中的幻觉
  • 多数幻觉内容倾向左翼,即使来自右翼来源
  • 高不确定性的生成阶段更易产生幻觉和偏移

大型语言模型(LLMs)被广泛用于回答政治信息类问题,尤其在选举相关场景中,事实错误与意识形态扭曲后果严重。我们构建了一个可复现的测量框架,将文档支撑下的问答幻觉视为意识形态漂移的诊断信号。基于涵盖左右翼及中间派来源的21,727篇美国政治新闻文章(来自QBias数据集),我们:(i) 为每篇文章生成特定问题;(ii) 从三个开源模型和一个专有模型获取文档支撑的回答;(iii) 通过参考对比检测句子级幻觉;(iv) 使用微调后的立场分类器对幻觉句进行意识形态归类;(v) 分析输出概率分布,探究词级别不确定性与幻觉及漂移的关系。结果显示,不同模型的幻觉率差异显著,集中在争议性话题;但来源意识形态对幻觉频率影响较小。然而,幻觉内容呈现显著左倾趋势:多数幻觉句被判定为左翼倾向,包括来自右翼来源的幻觉。词级概率分析表明,幻觉多出现在高熵生成情境,部分模型中不确定性还能预测左倾漂移,支持‘不确定性导致猜测’机制。研究对审计AI驱动的政治信息传播及选举相关部署的安全设计具有启示。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to answer questions about political information, including in election-adjacent information settings where factual errors and ideological distortions are high-stakes. We present a reproducible measurement framework that treats hallucinations, unsupported statements in document-grounded QA, as diagnostic signals of ideological drift. Using 21,727 expert-labeled U.S. political news articles from QBias spanning left, center, and right sources, we (i) generate an article-specific question, (ii) elicit document-grounded answers from three open-weight LLMs and one proprietary model, (iii) detect sentence-level hallucinations via reference-based comparison, (iv) classify the ideological valence of hallucinated sentences with a fine-tuned stance classifier, and (v) probe output logits to relate token-level uncertainty to hallucination and drift. Hallucination rates vary substantially across models and concentrate in contentious topics, while source-ideology differences in hallucination frequency are modest. In contrast, hallucination content exhibits robust leftward drift: a majority of hallucinated sentences are classified as left-leaning, including among hallucinations generated from right-leaning sources. Logit-level analysis shows hallucinations arise in high-entropy generation contexts, and in some models uncertainty also predicts leftward drift, consistent with an "uncertainty to guessing" mechanism. We discuss implications for auditing AI-mediated political information and for designing safeguards in election-relevant deployments.

大模型幻觉检测意识形态偏移政治问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。