用强化学习优化器提升模型对分布外数据的长期检测能力
Theoretical Grounding of Out-Of-Distribution Detection With Reinforcement Learning Optimizer
- 引入强化学习修正项,让优化过程主动降低未来语义分布外误报率
- 在动态环境中,相比传统方法,显著减少语义漂移下的假阳性
- 适合需要长期稳定检测能力的开放世界应用,如自动驾驶
在动态开放世界中进行分布外(OOD)检测,要求模型持续适应不断变化的数据分布,同时能泛化到协变量偏移输入,并拒绝语义偏移的分布外样本。现有大多数方法仅优化当前步骤目标,未显式考虑部署后环境变化对未来OOD行为的影响。本文基于强化学习(RL)构建动态OOD检测的理论基础,提出一种新的增强型优化器,在标准梯度下降(GD)基础上加入由强化学习引导的校正项,显式鼓励减少随时间推移的语义分布外假阳性率。该方法在面向未来域泛化和语义分布外拒绝任务上均表现更优。我们分析了时间误差分解,包括模型变化与环境变化带来的泛化误差,并建立了比较GD与强化学习引导优化器泛化误差的新理论框架。
原文摘要 · Abstract (English)
Out-of-distribution (OOD) detection in dynamic open-world environments requires a model to continually adapt to evolving data distributions while generalizing to covariate-shifted inputs and rejecting semantic-shifted OOD examples. Most existing OOD detection methods optimize only the current-step objective and do not explicitly account for how post-deployment environment changes affect future OOD behavior. In this paper, we establish a theoretical grounding for dynamic OOD detection using a reinforcement learning (RL)-guided optimizer that explicitly favors updates that reduce the semantic OOD false positive rate over time. We develop a novel augmented optimizer that uses an RL-guided correction term on top of standard gradient descent (GD) and show its improvement over both future-domain generalization and semantic-OOD rejection. We analyze temporal error decomposition in terms of model-change and environment-change generalization errors and develop a new theoretical framework for comparing the generalization errors under both GD and RL-guided optimizers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。