AI政治中立不可能实现,但可通过八种方法近似达成。
Political Neutrality in AI Is Impossible- But Here Is How to Approximate It
- 用程度化思路替代绝对中立,提出可操作的近似方法
- 在大模型输出层面验证框架,证明其可评估性
- 适合关注AI伦理与公平性的研究者和开发者
AI系统常表现出政治偏见,影响用户观点与决策。尽管政治中立(即无偏见)被视为公平与安全的理想目标,但本文指出其既不可行也非普遍可取,因中立具有主观性,且训练数据、算法与用户互动本身蕴含偏见。受约瑟夫·拉兹哲学思想启发——‘中立可以是程度问题’(Raz, 1986),我们主张追求一定程度的中立仍至关重要,以促进平衡的AI交互并减少用户操纵。因此,本文采用‘近似政治中立’概念,将焦点从不可达的绝对标准转向可实现的实践代理。提出八种在三个认知层级上实现近似中立的技术,分析其权衡与实施策略,并通过两个具体应用展示其实用性。最后,在大语言模型(LLMs)输出层面上评估该框架,演示其评估路径。本研究旨在推动对AI政治中立的深入讨论,促进负责任、对齐的语言模型发展。
原文摘要 · Abstract (English)
AI systems often exhibit political bias, influencing users' opinions and decisions. While political neutrality-defined as the absence of bias-is often seen as an ideal solution for fairness and safety, this position paper argues that true political neutrality is neither feasible nor universally desirable due to its subjective nature and the biases inherent in AI training data, algorithms, and user interactions. However, inspired by Joseph Raz's philosophical insight that "neutrality [...] can be a matter of degree" (Raz, 1986), we argue that striving for some neutrality remains essential for promoting balanced AI interactions and mitigating user manipulation. Therefore, we use the term "approximation" of political neutrality to shift the focus from unattainable absolutes to achievable, practical proxies. We propose eight techniques for approximating neutrality across three levels of conceptualizing AI, examining their trade-offs and implementation strategies. In addition, we explore two concrete applications of these approximations to illustrate their practicality. Finally, we assess our framework on current large language models (LLMs) at the output level, providing a demonstration of how it can be evaluated. This work seeks to advance nuanced discussions of political neutrality in AI and promote the development of responsible, aligned language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。