用偏好学习方法自动识别社会中不同群体的价值系统。
Learning the Value Systems of Societies with Preference-based Multi-objective Reinforcement Learning
- 基于聚类与偏好多目标强化学习,联合建模社会价值体系
- 在两个含人类价值观的MDP任务中优于现有算法
- 适合研究群体价值差异与可解释性价值对齐的学者
具备价值意识的AI应能识别人类价值观,并适应不同用户的價值體系(價值導向偏好)。这需要对價值進行操作化,但易導致錯誤定義。價值具有社會性,需在多用戶間保持一致,而價值體系雖多樣,卻在群體中呈現模式。在序列決策中,已有研究針對不同目標或價值進行個性化,但通常依賴手動設計特徵,缺乏價值導向可解釋性或對多樣用戶偏好的適應能力。本文提出基於聚類與偏好多目標強化學習(PbMORL)的算法,用於學習馬爾可夫決策過程(MDPs)中社會代理者群體的價值對齊模型與價值體系。我們聯合學習社會衍生的價值對齊模型(基礎)與一組簡潔代表不同用戶群體(聚類)的價值體系。每個聚類包含一個代表成員價值偏好之價值體系,以及一個近似帕累托最優的策略,反映與該價值體系一致的行為。我們在兩個含人類價值的MDP上評估該方法,並與現有先進的PbMORL算法及基線比較。
原文摘要 · Abstract (English)
Value-aware AI should recognise human values and adapt to the value systems (value-based preferences) of different users. This requires operationalization of values, which can be prone to misspecification. The social nature of values demands their representation to adhere to multiple users while value systems are diverse, yet exhibit patterns among groups. In sequential decision making, efforts have been made towards personalization for different goals or values from demonstrations of diverse agents. However, these approaches demand manually designed features or lack value-based interpretability and/or adaptability to diverse user preferences. We propose algorithms for learning models of value alignment and value systems for a society of agents in Markov Decision Processes (MDPs), based on clustering and preference-based multi-objective reinforcement learning (PbMORL). We jointly learn socially-derived value alignment models (groundings) and a set of value systems that concisely represent different groups of users (clusters) in a society. Each cluster consists of a value system representing the value-based preferences of its members and an approximately Pareto-optimal policy that reflects behaviours aligned with this value system. We evaluate our method against a state-of-the-art PbMORL algorithm and baselines on two MDPs with human values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。