通过观察与示范自动学习智能体的价值体系,实现跨价值系统的协同决策。
Learning the Value Systems of Agents with Preference-based and Inverse Reinforcement Learning
- 基于偏好与逆强化学习,从人类示范中推断价值权重
- 构建多目标马尔可夫决策过程模型,支持动态价值对齐
- 适用于需道德对齐的AI代理协作场景,如数字助手协商
协议技术指自主软件代理在开放系统中相互交互以达成双方接受协议的计算机系统。随着近年来人工智能的发展,这些协议要被各方接受,必须与伦理原则和道德价值观保持一致。然而,这极为困难,尤其因为不同人类用户(及其代理)可能持有不同的价值体系,即对各项道德价值的重要性判断不同。此外,在特定情境下精确计算定义某个价值也十分困难。基于人工设计规范(如价值调查)的方法受限于规模,因需要大量人工干预。本文提出一种新方法,可从观测数据和人类示范中自动学习价值体系。我们提出了价值体系学习的形式化模型,并将其应用于序贯决策领域,基于多目标马尔可夫决策过程,设计了专门的偏好型与逆强化学习算法,用于推断价值基础函数与价值体系。该方法通过两个模拟应用场景进行展示与评估。
原文摘要 · Abstract (English)
Agreement Technologies refer to open computer systems in which autonomous software agents interact with one another, typically on behalf of humans, in order to come to mutually acceptable agreements. With the advance of AI systems in recent years, it has become apparent that such agreements, in order to be acceptable to the involved parties, must remain aligned with ethical principles and moral values. However, this is notoriously difficult to ensure, especially as different human users (and their software agents) may hold different value systems, i.e. they may differently weigh the importance of individual moral values. Furthermore, it is often hard to specify the precise meaning of a value in a particular context in a computational manner. Methods to estimate value systems based on human-engineered specifications, e.g. based on value surveys, are limited in scale due to the need for intense human moderation. In this article, we propose a novel method to automatically \emph{learn} value systems from observations and human demonstrations. In particular, we propose a formal model of the \emph{value system learning} problem, its instantiation to sequential decision-making domains based on multi-objective Markov decision processes, as well as tailored preference-based and inverse reinforcement learning algorithms to infer value grounding functions and value systems. The approach is illustrated and evaluated by two simulated use cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。