arXiv:2607.11783cs.CL2026-07

温度影响大模型在检索增强生成中传递意识形态的强度

How Temperature Shapes Ideological Discourse in Retrieval-Augmented Generation?

论文配图:How Temperature Shapes Ideological Discourse in Retrieval-Augmented Generation?
图 1 · 摘自论文原文
  • 用语义分析法识别1117篇新冠治疗文章中的三种意识形态
  • 中等温度下生成内容与意识形态参考文本匹配度最高
  • 适合关注AI偏见、提示工程与可控生成的研究者

检索增强生成(RAG)被广泛用于减少大语言模型(LLM)的幻觉并增强其事实准确性。然而,检索内容中的意识形态偏见对输出的影响尚未被充分研究。本研究通过分析包含1,117篇新冠治疗文章的语料库,采用词汇多维分析(LMDA)识别出三种意识形态话语。该语料库作为RAG的外部知识源,评估多个LLM在不同采样温度下回答意识形态相关问题的表现。基于生成文本与意识形态参考文本的语义和词汇相似性进行评估。结果表明,RAG框架容易将意识形态话语传递至生成内容中,且采样温度显著影响传递强度:中等温度时,模型在随机性与检索依据间取得平衡,话语对齐度最高;低温时,过于确定的采样抑制了话语传递。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) has been increasingly adopted to reduce hallucinations and strengthen the factual grounding of large language models (LLMs). While robustness to errors in the retrieval process has been explored, the impact of ideological bias on LLM outputs has been overlooked. For instance, if the retrieved material contains ideological positions, the RAG may transmit, amplify, or suppress such ideological discourses in its outputs. In this study, we address this issue by examining the influence of the RAG framework, comprising ideological discourses, in LLM-generated answers. To this end, we applied Lexical Multidimensional Analysis (LMDA) on a corpus of 1,117 COVID-19 treatment articles, identifying three ideological discourses. This corpus is then used as the external knowledge source for the RAG. We assessed several LLMs by having the models answer ideological questions at different sampling temperatures. The generated texts were assessed semantically and lexically based on their similarities with ideological reference texts. Our findings show that the RAG framework is prone to transferring ideological discourses into LLM responses, with sampling temperature having a measurable impact on the strength of this transfer. Discoursive alignment between generated answers and the reference text is highest at moderate temperatures, where models balance stochasticity with retrieval grounding, and drops at low temperatures, indicating that overly deterministic sampling suppresses discourse transfer.

RAG意识形态温度控制大模型偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。