arXiv:2605.19215cs.AI2026-05

区分不确定性类型,发现波动性促探索,随机性抑制探索。

Not all uncertainty is alike: volatility, stochasticity, and exploration

论文配图:Not all uncertainty is alike: volatility, stochasticity, and exploration
图 1 · 摘自论文原文
  • 用高斯状态空间模型区分环境波动与观测噪声
  • 波动性增加探索,随机性减少探索,结果相反
  • 新算法CAUSE提升异质噪声环境下的学习效率

生物与人工智能中的自适应决策需权衡已知收益的利用与未知选项的探索。以往研究将不同来源的不确定性视为等同,但本文区分了随时间漂移的潜在奖励状态(波动性)与通过噪声观测带来的不确定性(随机性)。两者均增加后验不确定性,却导致最优探索方向相反:波动性增强探索,随机性抑制探索。通过将吉廷斯指数框架扩展至具有潜在动态的高斯状态空间老虎机问题,我们正式建立了这一不对称性。进一步推导出因果感知的不确定性敏感探索(CAUSE),一种基于控制即推理的闭式探索奖励,继承相同单调性。CAUSE在具有异质噪声结构的环境中优于标准探索策略,且超越仅适用于静止设置的吉廷斯每臂策略。学习与探索受同一噪声推断不对称性调控,该框架预测病理性噪声推断会导致探索反转而非简单减弱,对精神障碍的计算解释具有启示。

原文摘要 · Abstract (English)

Adaptive decision-making in biological and artificial intelligence requires balancing the exploitation of known outcomes with the exploration of uncertain alternatives. Although prior work suggests that uncertainty generally promotes exploration, it has typically treated distinct sources of environmental uncertainty as equivalent. We consider environments with latent reward states that drift over time (volatility) and are observed through noisy outcomes (stochasticity). Both increase posterior uncertainty, yet we show they drive optimal exploration in opposite directions: volatility enhances it, stochasticity suppresses it. We establish this asymmetry formally by extending the Gittins index framework to Gaussian state-space bandits with latent dynamics. We further derive Cause-Aware Uncertainty-Sensitive Exploration (CAUSE), a closed-form exploration bonus obtained via control-as-inference that inherits the same monotonicities. CAUSE outperforms standard exploration strategies in environments with heterogeneous noise structure, and also improves on a Gittins-per-arm policy whose rested-bandit optimality does not transfer to restless settings. Learning and exploration are governed by the same noise-inference asymmetry, and the framework predicts that pathological noise inference produces \emph{reversed} rather than merely impaired exploration, with implications for computational accounts of psychiatric conditions.

强化学习不确定性探索策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。