arXiv:2510.21638cs.LGcs.AI2025-10

用极简统计量实现快速可靠的强化学习异常检测

DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection

  • 仅用每回合均值和核相似度捕捉全局与局部偏差
  • 计算量减少600倍,准确率比主流方法高5个百分点
  • 适合对实时性与可靠性要求高的安全关键场景

在安全关键场景中部署强化学习(RL)受限于分布偏移下的脆弱性。本文研究了强化学习时间序列的分布外(OOD)检测问题,提出DEEDEE——一种基于双统计量的检测器,重新审视了以往依赖复杂表示的流程,提供了轻量级替代方案。DEEDEE仅使用每回合均值和与训练样本摘要的RBF核相似度,有效捕捉全局与局部偏离特征。尽管结构简单,其性能在标准RL OOD基准上达到或超过当前主流检测器,在计算量(FLOPs/运行时间)上实现600倍降低,平均准确率较强基线提升5%绝对值。概念上,结果表明多种异常类型常通过少量低阶统计量反映在RL轨迹中,提示复杂环境中可建立紧凑的分布外检测基础。

原文摘要 · Abstract (English)

Deploying reinforcement learning (RL) in safety-critical settings is constrained by brittleness under distribution shift. We study out-of-distribution (OOD) detection for RL time series and introduce DEEDEE, a two-statistic detector that revisits representation-heavy pipelines with a minimal alternative. DEEDEE uses only an episodewise mean and an RBF kernel similarity to a training summary, capturing complementary global and local deviations. Despite its simplicity, DEEDEE matches or surpasses contemporary detectors across standard RL OOD suites, delivering a 600-fold reduction in compute (FLOPs / wall-time) and an average 5% absolute accuracy gain over strong baselines. Conceptually, our results indicate that diverse anomaly types often imprint on RL trajectories through a small set of low-order statistics, suggesting a compact foundation for OOD detection in complex environments.

强化学习异常检测高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。