系统提示位置决定模型偏见,隐藏配置可能引发不可控歧视。
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
- 通过对比用户与系统提示中人口信息的处理差异,发现位置影响输出结果。
- 在6个商用大模型、50个人口群体中均观测到显著偏差,涉及代表性和决策差异。
- 适合关注AI公平性、部署安全性的开发者与监管者阅读。
大型语言模型(LLMs)中的系统提示是预设指令,优先于用户输入决定模型行为。尽管模型提供商设定基础提示,部署方和第三方开发者可附加提示,但这些叠加层对终端用户完全透明。随着提示结构复杂化,可能引入未被察觉的副作用。本文研究信息在不同指令中的位置如何影响模型行为,对比了6个商用大模型在50个人口群体中对人口信息的处理方式。结果表明存在显著偏见,表现为用户代表性差异与决策场景不公。由于这些偏差源于不可见且不透明的系统级配置,可能造成表征性、分配性及其他潜在危害,超出用户检测或纠正能力。研究呼吁将系统提示分析纳入AI审计流程,尤其当可定制系统提示日益普及时。
原文摘要 · Abstract (English)
System prompts in Large Language Models (LLMs) are predefined directives that guide model behaviour, taking precedence over user inputs in text processing and generation. LLM deployers increasingly use them to ensure consistent responses across contexts. While model providers set a foundation of system prompts, deployers and third-party developers can append additional prompts without visibility into others' additions, while this layered implementation remains entirely hidden from end-users. As system prompts become more complex, they can directly or indirectly introduce unaccounted for side effects. This lack of transparency raises fundamental questions about how the position of information in different directives shapes model outputs. As such, this work examines how the placement of information affects model behaviour. To this end, we compare how models process demographic information in system versus user prompts across six commercially available LLMs and 50 demographic groups. Our analysis reveals significant biases, manifesting in differences in user representation and decision-making scenarios. Since these variations stem from inaccessible and opaque system-level configurations, they risk representational, allocative and potential other biases and downstream harms beyond the user's ability to detect or correct. Our findings draw attention to these critical issues, which have the potential to perpetuate harms if left unexamined. Further, we argue that system prompt analysis must be incorporated into AI auditing processes, particularly as customisable system prompts become increasingly prevalent in commercial AI deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。