arXiv:2608.10327cs.AI2026-08

94篇论文分析揭示AI对人类价值的隐含定义及其潜在问题

Toward a Theory of Value in AI Alignment

论文配图:Toward a Theory of Value in AI Alignment
图 1 · 摘自论文原文
  • 通过分析94篇论文,发现多数研究将价值简化为偏好
  • 依赖合成数据和自动评分,可能忽略文化语境中的复杂价值
  • 呼吁明确哲学立场,拓展非主流价值表达路径

大型语言模型和多模态基础模型的普及带来了从有毒言论、幻觉到未经授权行动等各类危害。在人工智能安全领域,这些现象常被归因于对齐问题,即模型与人类价值观不一致。研究人员虽开展应用与理论对齐工作,却很少明确定义“人类价值观”具体所指。本研究对94篇对齐论文进行标注,揭示其隐含的价值观理论:多数未明确定义价值,转而以偏好代替,易将复杂文化情境下的概念简化为二元选择。随着研究者逐步放弃人工标注,转向合成数据与自动评分系统进行模型训练与评估,可能封闭其他挑战与实践价值观的途径。本文旨在使人工智能对齐背后的哲学预设显性化,推动更具体且未被充分探讨的观点进入关于人工智能如何回应人类价值观的讨论。

原文摘要 · Abstract (English)

Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as the alignment problem, or of models being misaligned with human values. Researchers have responded by pursuing applied and theoretical AI value alignment efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? We annotated 94 value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on preferences as a stand in that runs the risk of reducing complex culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and autorater approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. In making AI value alignments philosophical commitments explicit, we seek to bring great specificity and under explored perspectives in the debate on whether and how AI can address human values.

AI对齐价值观伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。