arXiv:2511.12378cs.AI2025-11

让智能体动态判断建议可靠性,自主决定何时求助。

Learning to Trust: Bayesian Adaptation to Varying Suggester Reliability in Sequential Decision Making

  • 用贝叶斯推理实时评估建议者可信度,动态调整依赖程度。
  • 在不同建议质量下均表现稳健,能适应可靠性变化。
  • 适合人机协作场景,尤其适用于不确定环境中的决策系统。

在部分可观测环境中,自主智能体在序列决策任务中可受益于外部动作建议,但这些建议的可靠性存在差异。现有方法通常假设建议者质量恒定且已知,限制了实际应用。本文提出一种动态学习并适应建议者可靠性变化的框架:首先,将建议者质量直接嵌入智能体的信念表示中,通过贝叶斯推理推断建议类型,从而自适应调整对建议的依赖;其次,引入显式的“请求”动作,使智能体能在关键时刻策略性地索取建议,权衡信息收益与获取成本。实验表明,该方法在不同建议质量下均表现稳健,能适应可靠性变化,并实现对建议请求的智能管理。本工作为应对不确定环境中的建议不确定性,奠定了自适应人机协作的基础。

原文摘要 · Abstract (English)

Autonomous agents operating in sequential decision-making tasks under uncertainty can benefit from external action suggestions, which provide valuable guidance but inherently vary in reliability. Existing methods for incorporating such advice typically assume static and known suggester quality parameters, limiting practical deployment. We introduce a framework that dynamically learns and adapts to varying suggester reliability in partially observable environments. First, we integrate suggester quality directly into the agent's belief representation, enabling agents to infer and adjust their reliance on suggestions through Bayesian inference over suggester types. Second, we introduce an explicit ``ask'' action allowing agents to strategically request suggestions at critical moments, balancing informational gains against acquisition costs. Experimental evaluation demonstrates robust performance across varying suggester qualities, adaptation to changing reliability, and strategic management of suggestion requests. This work provides a foundation for adaptive human-agent collaboration by addressing suggestion uncertainty in uncertain environments.

强化学习贝叶斯推理人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。