arXiv:2602.03939eess.SYcs.LG2026-02

通过信息导向目标,让智能体在探索中更快识别环境上下文并提升收益。

C-IDS: Solving Contextual POMDP via Information-Directed Objective

  • 设计信息导向目标,将上下文与观测的互信息纳入优化
  • 算法在连续光暗环境中实现更快上下文识别和更高累计回报
  • 理论证明温度参数是信息率上界,支持子线性贝叶斯后悔

我们研究了上下文部分可观测马尔可夫决策过程(CPOMDP)中的策略合成问题,其中环境由未知潜在上下文控制,导致不同的POMDP动态。目标是设计一种策略,既能最大化累计回报,又能主动减少对潜在上下文的不确定性。我们提出一种信息导向目标,将奖励最大化与潜在上下文和智能体观测之间的互信息相结合。我们开发了C-IDS算法,用于合成最大化该信息导向目标的策略。我们证明该目标可解释为线性信息率的拉格朗日松弛,并表明温度参数是信息率的上界。基于此特性,我们建立了在K个阶段内的子线性贝叶斯后悔界。我们在连续光暗环境中评估了该方法,结果表明其持续优于将未知上下文视为隐状态变量的标准POMDP求解器,实现了更快的上下文识别和更高的回报。

原文摘要 · Abstract (English)

We study the policy synthesis problem in contextual partially observable Markov decision processes (CPOMDPs), where the environment is governed by an unknown latent context that induces distinct POMDP dynamics. Our goal is to design a policy that simultaneously maximizes cumulative return and actively reduces uncertainty about the underlying context. We introduce an information-directed objective that augments reward maximization with mutual information between the latent context and the agent's observations. We develop the C-IDS algorithm to synthesize policies that maximize the information-directed objective. We show that the objective can be interpreted as a Lagrangian relaxation of the linear information ratio and prove that the temperature parameter is an upper bound on the information ratio. Based on this characterization, we establish a sublinear Bayesian regret bound over K episodes. We evaluate our approach on a continuous Light-Dark environment and show that it consistently outperforms standard POMDP solvers that treat the unknown context as a latent state variable, achieving faster context identification and higher returns.

强化学习上下文感知信息率策略合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。