构建生物经济对齐的多智能体安全测试基准,揭示智能体行为的关键风险。
From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks
- 基于生物稳态与经济递减收益设计多目标多智能体环境
- 8个基准场景揭示资源耗尽、目标失衡等核心安全问题
- 适合研究AI对齐与可持续智能体系统的学者参考
构建安全、对齐的智能体系统需要全面的实证测试,但现有基准普遍忽视与生物学和经济学相关的关键主题——这两门经验证的学科深刻描述了人类需求与偏好。为此,本文提出一组以生物和经济原理为动机的多目标、多智能体对齐基准,强调有限生物性目标的稳态维持、无限工具性与商业目标的递减收益、可持续性原则及资源共享。共实现8个主要基准环境,用于揭示智能体系统中的关键陷阱与挑战,如无限制最大化稳态目标、过度优化单一目标而牺牲其他目标、忽视安全约束,或耗尽共享资源。
原文摘要 · Abstract (English)
Developing safe, aligned agentic AI systems requires comprehensive empirical testing, yet many existing benchmarks neglect crucial themes aligned with biology and economics, both time-tested fundamental sciences describing our needs and preferences. To address this gap, the present work focuses on introducing biologically and economically motivated themes that have been neglected in current mainstream discussions on AI safety - namely a set of multi-objective, multi-agent alignment benchmarks that emphasize homeostasis for bounded and biological objectives, diminishing returns for unbounded, instrumental, and business objectives, sustainability principle, and resource sharing. Eight main benchmark environments have been implemented on the above themes, to illustrate key pitfalls and challenges in agentic AI-s, such as unboundedly maximizing a homeostatic objective, over-optimizing one objective at the expense of others, neglecting safety constraints, or depleting shared resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。