arXiv:2510.14315cs.LGstat.ML2025-10被引 1

智能体自主决定何时测量状态,兼顾成本与长期效果。

Active Measuring in Reinforcement Learning With Delayed Negative Effects

  • 引入可主动观测的MDP模型,决策是否测量隐状态
  • 测量虽延迟有害但能提升采样效率与策略价值
  • 适用于需权衡评估成本的数字健康等场景

在强化学习中,状态测量可能代价高昂且对后续环境产生延迟负面影响。本文提出主动可观测马尔可夫决策过程(AOMDP),其中智能体不仅选择控制动作,还决定是否测量潜在状态。测量动作可揭示真实状态,但可能带来延迟负效应。我们证明,尽管存在此类代价,减少不确定性仍可显著提升样本效率并提高最优策略价值。将AOMDP建模为周期性部分可观测MDP,提出基于信念状态的在线强化学习算法。为近似信念状态,设计一种序列蒙特卡洛方法,联合估计未知静态环境参数和未观测隐状态。在数字健康应用中评估:智能体决定何时推送数字干预、何时通过问卷评估用户健康状况。

原文摘要 · Abstract (English)

Measuring states in reinforcement learning (RL) can be costly in real-world settings and may negatively influence future outcomes. We introduce the Actively Observable Markov Decision Process (AOMDP), where an agent not only selects control actions but also decides whether to measure the latent state. The measurement action reveals the true latent state but may have a negative delayed effect on the environment. We show that this reduced uncertainty may provably improve sample efficiency and increase the value of the optimal policy despite these costs. We formulate an AOMDP as a periodic partially observable MDP and propose an online RL algorithm based on belief states. To approximate the belief states, we further propose a sequential Monte Carlo method to jointly approximate the posterior of unknown static environment parameters and unobserved latent states. We evaluate the proposed algorithm in a digital health application, where the agent decides when to deliver digital interventions and when to assess users' health status through surveys.

强化学习状态测量数字健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。