用信息导向采样解决智能体与人类偏好对齐的探索难题
Aligning AI Agents via Information-Directed Sampling
- 提出带人类偏好反馈的多臂老虎机框架,需平衡环境与人类偏好探索
- 信息导向采样在后悔值上显著优于传统方法如汤普森采样
- 适合研究人机对齐、强化学习中主动学习的学者参考
AI系统的惊人成就凸显了人工智能对齐问题:如何使‘超智能’智能体的行为与人类利益一致。现有方法大多局限于短期视野或孤立地学习人类反馈,且假设智能体已完全识别环境。为此,本文将经典多臂老虎机问题扩展为一类带对齐约束的贝叶斯老虎机问题,其中智能体需通过与未知环境和人类的交互最大化长期期望回报。动作奖励取决于可观测结果与人类偏好,而查询人类有成本。因此,有效智能体需智能权衡探索(环境与人类)与利用。我们在一个类beta-Bernoulli的简化问题中理论与实证分析该权衡,发现当前普遍采用的探索算法甚至被推崇的汤普森采样均无法提供可接受解,而信息导向采样(Information-Directed Sampling)展现出更优的后悔表现。
原文摘要 · Abstract (English)
The staggering feats of AI systems have brought to attention the topic of AI Alignment: aligning a "superintelligent" AI agent's actions with humanity's interests. Many existing frameworks/algorithms in alignment study the problem on a myopic horizon or study learning from human feedback in isolation, relying on the contrived assumption that the agent has already perfectly identified the environment. As a starting point to address these limitations, we define a class of bandit alignment problems as an extension of classic multi-armed bandit problems. A bandit alignment problem involves an agent tasked with maximizing long-run expected reward by interacting with an environment and a human, both involving details/preferences initially unknown to the agent. The reward of actions in the environment depends on both observed outcomes and human preferences. Furthermore, costs are associated with querying the human to learn preferences. Therefore, an effective agent ought to intelligently trade-off exploration (of the environment and human) and exploitation. We study these trade-offs theoretically and empirically in a toy bandit alignment problem which resembles the beta-Bernoulli bandit. We demonstrate while naive exploration algorithms which reflect current practices and even touted algorithms such as Thompson sampling both fail to provide acceptable solutions to this problem, information-directed sampling achieves favorable regret.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。