arXiv:2506.01080cs.AIcs.CY2025-06被引 22

多智能体系统需动态对齐人类价值,否则可能集体偏离目标。

The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process

  • 将对齐视为社会互动中的动态过程,而非静态目标
  • 多智能体协作易导致集体偏离人类价值观
  • 适合关注AI社会影响与系统设计的研究者

本文指出,多智能体系统(MAS)中的AI对齐应被视为一种动态且依赖社会环境的过程,受合作、协同或竞争关系影响。随着多智能体系统在现实应用中日益普及,其复杂的交互机制会重塑智能体的目标追求方式。当智能体相互协作以达成个体与集体目标时,可能无意中使部分或全部智能体偏离人类价值观或用户偏好。借鉴社会科学,本文分析社会结构如何削弱或瓦解群体与个体价值。因此,呼吁人工智能界将人类偏好与客观对齐视为相互依存的概念,而非孤立问题。最后强调,亟需构建仿真环境、基准测试与评估框架,以在动态多智能体交互场景中提前评估对齐状态,防止系统复杂性失控。

原文摘要 · Abstract (English)

This position paper states that AI Alignment in Multi-Agent Systems (MAS) should be considered a dynamic and interaction-dependent process that heavily depends on the social environment where agents are deployed, either collaborative, cooperative, or competitive. While AI alignment with human values and preferences remains a core challenge, the growing prevalence of MAS in real-world applications introduces a new dynamic that reshapes how agents pursue goals and interact to accomplish various tasks. As agents engage with one another, they must coordinate to accomplish both individual and collective goals. However, this complex social organization may unintentionally misalign some or all of these agents with human values or user preferences. Drawing on social sciences, we analyze how social structure can deter or shatter group and individual values. Based on these analyses, we call on the AI community to treat human, preferential, and objective alignment as an interdependent concept, rather than isolated problems. Finally, we emphasize the urgent need for simulation environments, benchmarks, and evaluation frameworks that allow researchers to assess alignment in these interactive multi-agent contexts before such dynamics grow too complex to control.

多智能体对齐社会性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。