AI应动态适应人类价值观变化,而非僵化对齐。
Rethinking How AI Embeds and Adapts to Human Values: Challenges and Opportunities
- 提出动态价值对齐框架,强调长期推理与适应性。
- 指出单一静态价值模型易引发意外风险,需多主体协同应对多元价值冲突。
- 适合关注AI伦理、人机协作及价值对齐研究者阅读。
「以人为本的AI」和「基于价值的决策」在学术界与工业界备受关注。然而,诸多关键问题仍待深入探讨,包括系统如何嵌入人类价值观、人类如何识别系统中的价值观,以及如何最小化伤害或意外后果。本文呼吁重新思考价值对齐的范式,主张其不应局限于静态、单一的价值观。强调AI系统应具备长期推理能力,并能适应价值的演化。此外,价值对齐需更丰富的理论支撑以涵盖人类价值观的全谱系。由于价值观常因个体或群体而异,多智能体系统为处理多元主义、冲突及跨智能体价值推理提供了合适框架。本文揭示了价值对齐面临的挑战,并指明了未来研究方向。同时,从设计方法到实际应用,广泛讨论了价值对齐的不同视角。
原文摘要 · Abstract (English)
The concepts of ``human-centered AI'' and ``value-based decision'' have gained significant attention in both research and industry. However, many critical aspects remain underexplored and require further investigation. In particular, there is a need to understand how systems incorporate human values, how humans can identify these values within systems, and how to minimize the risks of harm or unintended consequences. In this paper, we highlight the need to rethink how we frame value alignment and assert that value alignment should move beyond static and singular conceptions of values. We argue that AI systems should implement long-term reasoning and remain adaptable to evolving values. Furthermore, value alignment requires more theories to address the full spectrum of human values. Since values often vary among individuals or groups, multi-agent systems provide the right framework for navigating pluralism, conflict, and inter-agent reasoning about values. We identify the challenges associated with value alignment and indicate directions for advancing value alignment research. In addition, we broadly discuss diverse perspectives of value alignment, from design methodologies to practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。