arXiv:2608.27910cs.AIcs.CL2026-08中稿 · EMNLP综述

用博弈论视角重新审视AI对齐难题,揭示现有方法的优劣与局限。

AI Alignment through a Game-theoretic Lens: A Survey

论文配图:AI Alignment through a Game-theoretic Lens: A Survey
图 1 · 摘自论文原文
  • 将人类偏好建模为多方动态博弈,捕捉复杂情境下的非传递性选择
  • 指出当前对齐方法在处理多元偏好和时间演化时存在理论短板
  • 适合关注AI伦理、多智能体系统设计的研究者阅读

随着大语言模型和日益强大的AI代理被部署于高风险场景,如何使其与复杂的现实人类价值观对齐已成为核心挑战。现有对齐方法虽能提升有用性、无害性和可控性,但难以捕捉情境依赖、非传递性且受多方动态互动影响的真实偏好。本文从博弈论视角综述AI对齐研究,围绕关键博弈要素梳理近期进展,聚焦三大挑战:偏好多样性、对齐优先级与时间动态。该框架厘清了当前方法在博弈论分析中的真实价值、适用边界及尚未解决的难题,为构建稳健、自适应且可验证的AI系统指明方向。

原文摘要 · Abstract (English)

As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability, often struggle to capture real-world preferences that are context-dependent, non-transitive, and shaped by dynamic multi-party interactions. This survey reviews AI alignment through a game-theoretic lens. Specifically, it organizes recent progress around key game-theoretic elements and synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics. This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.

AI对齐博弈论多智能体伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。