让智能体通过观察他人行为,自动识别高手并高效学习。
Exploiting Expertise of Non-Expert and Diverse Agents in Social Bandit Learning: A Free Energy Approach
- 基于自由能原理,在无奖励信息下评估他人能力
- 在多种场景中显著提升学习效果,保持对数级后悔值
- 适合存在非专家但有用信息的复杂协作环境
个性化AI服务涉及大量独立强化学习智能体。现有算法多关注个体学习,忽视人类和动物普遍存在的社会学习能力。本文研究一种社交老虎机学习场景:社交智能体仅观察他人行为,无法获知其奖励。各智能体独立追求自身策略,无教学动机。我们提出一种基于自由能的社交老虎机学习算法,直接在策略空间上运作,无需任何先验知识或社会规范即可评估他人专业水平。该算法融合自身环境经验与他人估计策略。理论证明算法收敛至最优策略。实证结果表明,该方法在各类场景中优于现有方法,能有效识别相关智能体,即使面对随机或低效智能体也能利用其行为信息。尤其在存在相关但非专家智能体时,显著提升学习性能,而多数现有方法在此失效。更重要的是,算法维持对数级后悔值。
原文摘要 · Abstract (English)
Personalized AI-based services involve a population of individual reinforcement learning agents. However, most reinforcement learning algorithms focus on harnessing individual learning and fail to leverage the social learning capabilities commonly exhibited by humans and animals. Social learning integrates individual experience with observing others' behavior, presenting opportunities for improved learning outcomes. In this study, we focus on a social bandit learning scenario where a social agent observes other agents' actions without knowledge of their rewards. The agents independently pursue their own policy without explicit motivation to teach each other. We propose a free energy-based social bandit learning algorithm over the policy space, where the social agent evaluates others' expertise levels without resorting to any oracle or social norms. Accordingly, the social agent integrates its direct experiences in the environment and others' estimated policies. The theoretical convergence of our algorithm to the optimal policy is proven. Empirical evaluations validate the superiority of our social learning method over alternative approaches in various scenarios. Our algorithm strategically identifies the relevant agents, even in the presence of random or suboptimal agents, and skillfully exploits their behavioral information. In addition to societies including expert agents, in the presence of relevant but non-expert agents, our algorithm significantly enhances individual learning performance, where most related methods fail. Importantly, it also maintains logarithmic regret.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。