检测大模型是否被刻意引导至特定意识形态
Don't Change My View: Ideological Bias Auditing in Large Language Models
- 通过分析提示词相关输出分布变化,识别意识形态偏移
- 无需访问模型内部,可对黑箱系统进行审计
- 适合独立机构事后审查大模型行为
随着大型语言模型(LLMs)在数百万用户产品中广泛应用,其输出可能影响个体信念,并累积形成公共舆论。若大模型的行为可被有意引导至特定意识形态立场(如政治或宗教观点),则系统控制者可能获得对公共话语的不成比例影响力。尽管目前尚不清楚大模型是否能可靠地被引导至一致的意识形态立场,以及此类引导能否有效防范,但首要步骤是开发检测此类引导行为的方法。本文将一种先前提出的统计方法适配到意识形态偏见审计的新场景中。该方法继承了原框架的模型无关设计,无需访问语言模型内部结构,而是通过分析与特定主题相关的提示词下模型输出的分布变化来识别潜在的意识形态引导。这一设计使其特别适用于对专有黑箱系统的审计。我们通过一系列实验验证了该方法的实用性及其在支持独立后置审计方面的潜力。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly embedded in products used by millions, their outputs may influence individual beliefs and, cumulatively, shape public opinion. If the behavior of LLMs can be intentionally steered toward specific ideological positions, such as political or religious views, then those who control these systems could gain disproportionate influence over public discourse. Although it remains an open question whether LLMs can reliably be guided toward coherent ideological stances and whether such steering can be effectively prevented, a crucial first step is to develop methods for detecting when such steering attempts occur. In this work, we adapt a previously proposed statistical method to the new context of ideological bias auditing. Our approach carries over the model-agnostic design of the original framework, which does not require access to the internals of the language model. Instead, it identifies potential ideological steering by analyzing distributional shifts in model outputs across prompts that are thematically related to a chosen topic. This design makes the method particularly suitable for auditing proprietary black-box systems. We validate our approach through a series of experiments, demonstrating its practical applicability and its potential to support independent post hoc audits of LLM behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。