arXiv:2509.13854cs.CYcs.AI2025-09被引 3

厘清人机价值对齐的内涵与实现路径

Understanding the Process of Human-AI Value Alignment

  • 通过分析172篇论文,提炼出六类核心研究主题
  • 提出价值对齐是持续互动的过程,需兼顾认知局限与伦理冲突
  • 适合关注AI伦理、人机协作的研究者参考

背景:计算机科学中常将价值对齐用于描述人工智能与人类的一致性,但该术语使用往往缺乏精确性。目标:本文通过系统文献综述,深入理解人工智能中的价值对齐问题,基于研究文献提出更精准的定义。方法:分析近多年发表的172篇价值对齐研究论文,采用主题分析法进行内容整合。结果:分析归纳出六个主题:价值对齐的驱动力与方法;价值对齐面临的挑战;价值对齐中的价值观;人类与AI的认知过程;人-代理协同;价值对齐系统的设计与开发。结论:基于文献语境,我们将价值对齐定义为人类与自主代理之间持续进行的过程,旨在表达并实施抽象价值,同时管理人类与AI的认知局限,并平衡不同群体间因价值观差异产生的伦理与政治矛盾。本研究揭示了该领域未来的研究挑战与机遇。

原文摘要 · Abstract (English)

Background: Value alignment in computer science research is often used to refer to the process of aligning artificial intelligence with humans, but the way the phrase is used often lacks precision. Objectives: In this paper, we conduct a systematic literature review to advance the understanding of value alignment in artificial intelligence by characterising the topic in the context of its research literature. We use this to suggest a more precise definition of the term. Methods: We analyse 172 value alignment research articles that have been published in recent years and synthesise their content using thematic analyses. Results: Our analysis leads to six themes: value alignment drivers & approaches; challenges in value alignment; values in value alignment; cognitive processes in humans and AI; human-agent teaming; and designing and developing value-aligned systems. Conclusions: By analysing these themes in the context of the literature we define value alignment as an ongoing process between humans and autonomous agents that aims to express and implement abstract values in diverse contexts, while managing the cognitive limits of both humans and AI agents and also balancing the conflicting ethical and political demands generated by the values in different groups. Our analysis gives rise to a set of research challenges and opportunities in the field of value alignment for future work.

AI伦理人机协作价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。