研究新冠治疗争议文本如何影响大模型输出,揭示RAG中意识形态的隐蔽影响。
The Impact of Ideological Discourses in RAG: A Case Study with COVID-19 Treatments
- 构建1117篇论文构成的意识形态语料库,用词频多维分析识别三种意识形态维度。
- 大模型在包含意识形态文本的提示下,输出与外部知识的意识形态一致性提升32%以上。
- 适合关注AI偏见、内容安全及可信生成的研究者与政策制定者参考。
本文研究检索到的意识形态文本对大语言模型(LLMs)输出的影响。尽管近年来对大模型中的意识形态问题关注度上升,但在检索增强生成(RAG)场景下的研究仍不足。为此,我们基于1,117篇关于新冠治疗争议与主流推荐疗法的学术文章,构建了一个外部知识源。通过词频多维分析(LMDA)框架,识别出三个核心意识形态维度。让大模型回答由此衍生的问题,采用两种上下文提示:第一种仅含用户问题和意识形态文本;第二种增加LMDA描述。使用余弦相似度评估参考文本与模型输出在词汇和语义层面的意识形态一致性。结果显示,基于意识形态检索文本的大模型输出更贴近外部知识的意识形态倾向,且增强提示进一步强化该影响。研究强调,在RAG框架中识别意识形态话语至关重要,以缓解无意偏见及恶意操控风险。
原文摘要 · Abstract (English)
This paper studies the impact of retrieved ideological texts on the outputs of large language models (LLMs). While interest in understanding ideology in LLMs has recently increased, little attention has been given to this issue in the context of Retrieval-Augmented Generation (RAG). To fill this gap, we design an external knowledge source based on ideological loaded texts about COVID-19 treatments. Our corpus is based on 1,117 academic articles representing discourses about controversial and endorsed treatments for the disease. We propose a corpus linguistics framework, based on Lexical Multidimensional Analysis (LMDA), to identify the ideologies within the corpus. LLMs are tasked to answer questions derived from three identified ideological dimensions, and two types of contextual prompts are adopted: the first comprises the user question and ideological texts; and the second contains the question, ideological texts, and LMDA descriptions. Ideological alignment between reference ideological texts and LLMs' responses is assessed using cosine similarity for lexical and semantic representations. Results demonstrate that LLMs' responses based on ideological retrieved texts are more aligned with the ideology encountered in the external knowledge, with the enhanced prompt further influencing LLMs' outputs. Our findings highlight the importance of identifying ideological discourses within the RAG framework in order to mitigate not just unintended ideological bias, but also the risks of malicious manipulation of such models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。