高频率决策下,传统RL失效,本文提出用分布优势提升性能
Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
- 发现高频率时动作收益分布会坍缩,导致传统方法失准
- 量化了分布坍缩速率,证明统计量衰减速度不同
- 引入概率化优势概念,构建新算法提升高频控制效果
当决策频率较高时,传统强化学习方法难以准确估计动作价值,导致性能不稳定且表现差。分布式强化学习(DRL)在类似场景下的表现尚不明确。本文证明DRL代理对决策频率敏感:随着频率增加,动作条件回报分布会坍缩至基础策略的回报分布。我们量化了该坍缩速率,并发现不同统计量以不同速率衰减。进一步地,我们定义了分布视角下的动作差距与优势,提出‘优越性’作为优势的随机推广——这是缓解高频价值型RL性能问题的核心对象。此外,我们构建了一种基于优越性的DRL算法。在期权交易领域的仿真中验证,正确建模优越性分布可显著提升高频决策下的控制器性能。
原文摘要 · Abstract (English)
When decisions are made at high frequency, traditional reinforcement learning (RL) methods struggle to accurately estimate action values. In turn, their performance is inconsistent and often poor. Whether the performance of distributional RL (DRL) agents suffers similarly, however, is unknown. In this work, we establish that DRL agents are sensitive to the decision frequency. We prove that action-conditioned return distributions collapse to their underlying policy's return distribution as the decision frequency increases. We quantify the rate of collapse of these return distributions and exhibit that their statistics collapse at different rates. Moreover, we define distributional perspectives on action gaps and advantages. In particular, we introduce the superiority as a probabilistic generalization of the advantage -- the core object of approaches to mitigating performance issues in high-frequency value-based RL. In addition, we build a superiority-based DRL algorithm. Through simulations in an option-trading domain, we validate that proper modeling of the superiority distribution produces improved controllers at high decision frequencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。