用多目标强化学习动态调整AI偏好,适配多样变化的用户需求。
Adaptive Alignment: Dynamic Preference Adjustments via Multi-Objective Reinforcement Learning for Pluralistic AI
- 通过后训练策略选择实现偏好动态调整
- 基于多目标强化学习适应复杂多变的用户需求
- 适合需持续响应多元价值的智能系统设计
拟态人工智能(Pluralistic AI)对齐研究致力于解决智能系统如何根据多样化的用户需求与价值观进行设计和部署。本文提出一种基于多目标强化学习(MORL)的动态对齐方法,通过后训练阶段的策略选择调整,实现AI对多样化且动态变化的用户偏好的适应。论文介绍了该框架的设计思路、预期优势与假设条件,并讨论了技术实现细节。同时,从社会技术系统视角探讨了采用逆向对齐方法的广泛影响。
原文摘要 · Abstract (English)
Emerging research in Pluralistic Artificial Intelligence (AI) alignment seeks to address how intelligent systems can be designed and deployed in accordance with diverse human needs and values. We contribute to this pursuit with a dynamic approach for aligning AI with diverse and shifting user preferences through Multi Objective Reinforcement Learning (MORL), via post-learning policy selection adjustment. In this paper, we introduce the proposed framework for this approach, outline its anticipated advantages and assumptions, and discuss technical details about the implementation. We also examine the broader implications of adopting a retroactive alignment approach through the sociotechnical systems perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。