arXiv:2510.08113cs.LGcs.AI2025-10

教智能体如何判断何时信任专家数据,提升决策效率。

Bayesian Decision Making around Experts

  • 用贝叶斯方法建模专家数据对学习者信念的影响。
  • 预训练专家数据可降低信息论后悔边界,降幅由互信息决定。
  • 提出动态选择信任专家或自身体验的智能策略。

复杂学习智能体越来越多地与人类操作员或已有训练好的智能体协同工作。然而,学习者如何最优地利用结构上不同于自身动作-结果经验的专家数据仍不明确。本文在贝叶斯多臂赌博机框架下研究两类场景:(i) 离线设置,学习者在交互前接收专家最优策略产生的结果数据集;(ii) 同步设置,学习者每步需决定是基于自身经验更新信念,还是基于专家同时达成的结果。我们形式化了专家数据对学习者后验的影响,并证明:在离线设置中,预训练专家结果能通过专家数据与最优动作间的互信息,收紧信息论后悔界限。对于同步设置,我们提出一种信息导向规则,即选择能最大化一步信息增益的数据源。最后,我们设计策略使学习者能推断何时可信专家、何时不可信,以防范专家失效或被破坏的情况。本框架量化了专家数据的价值,为智能体提供实用且信息论基础的决策算法。

原文摘要 · Abstract (English)

Complex learning agents are increasingly deployed alongside existing experts, such as human operators or previously trained agents. However, it remains unclear how should learners optimally incorporate certain forms of expert data, which may differ in structure from the learner's own action-outcome experiences. We study this problem in the context of Bayesian multi-armed bandits, considering: (i) offline settings, where the learner receives a dataset of outcomes from the expert's optimal policy before interaction, and (ii) simultaneous settings, where the learner must choose at each step whether to update its beliefs based on its own experience, or based on the outcome simultaneously achieved by an expert. We formalize how expert data influences the learner's posterior, and prove that pretraining on expert outcomes tightens information-theoretic regret bounds by the mutual information between the expert data and the optimal action. For the simultaneous setting, we propose an information-directed rule where the learner processes the data source that maximizes their one-step information gain about the optimal action. Finally, we propose strategies for how the learner can infer when to trust the expert and when not to, safeguarding the learner for the cases where the expert is ineffective or compromised. By quantifying the value of expert data, our framework provides practical, information-theoretic algorithms for agents to intelligently decide when to learn from others.

贝叶斯决策专家数据多臂赌博机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。