arXiv:2501.19401cs.LGstat.ML2025-01

无需先验知识,用检测机制让旧算法适应动态环境。

DAL: A Practical Prior-Free Black-Box Framework for Piecewise Stationary Bandits

  • 用变化检测器增强现有算法,实现黑盒适配
  • 在多种真实与合成数据上表现优于当前最优方法
  • 适合需要快速部署的动态决策场景

我们提出一种实用的黑盒框架 Detection Augmented Learning (DAL),用于解决未知非平稳性的分段平稳老虎机问题。DAL 接受任何具有最优后悔率的平稳老虎机算法作为输入,并通过引入变化检测器,使其适用于所有常见的老虎机变体。大量实验表明,DAL 在多样化的非平稳场景中(包括合成基准和真实世界数据集)始终优于所有现有先进方法,彰显其通用性与可扩展性。我们提供了关于 DAL 强大实证性能的理论分析,并辅以全面的实证验证。

原文摘要 · Abstract (English)

We introduce a practical, black-box framework termed Detection Augmented Learning (DAL) for the problem of piecewise stationary bandits without knowledge of the underlying non-stationarity. DAL accepts any stationary bandit algorithm with order-optimal regret as input and augments it with a change detector, enabling applicability to all common bandit variants. Extensive experimentation demonstrates that DAL consistently surpasses all state-of-the-art methods across diverse non-stationary scenarios, including synthetic benchmarks and real-world datasets, underscoring its versatility and scalability. We provide theoretical insights into DAL's strong empirical performance, complemented by thorough empirical validation.

强化学习在线学习动态环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。