arXiv:2608.04305cs.LGq-fin.RM2026-08

为风险敏感强化学习设计自适应训练控制器,提升金融场景下策略的稳定性与收益质量。

Adaptive Finite-Budget Training for CVaR Risk-Aware Q-Learning

  • 通过动态调整内层步长和采样策略,优化有限预算下的训练过程
  • 使贝尔曼残差降低85%,在多种风险水平下保持稳定性能
  • 适合关注金融交易中风险控制与鲁棒性的研究者和实践者

风险敏感Q-learning(RaQL)提供了一种无模型、双时尺度的动态风险目标估计器,但其有限预算行为仍不稳定:固定内层超参数会导致值估计波动、持续的贝尔曼残差以及低效的样本重用。本文提出一种针对条件风险价值(CVaR)RaQL的自适应训练控制器,并在每日比特币交易任务上进行评估。该控制器保留原始CVaR估计器与贝尔曼不动点,通过六项协同机制重构训练流程:单元级内层步长调节、外层学习率匹配的衰减同步、对类似VaR的内变量进行早期短时修正、先覆盖后贪婪的样本分配规则、成熟内估计的渐进后缀聚合,以及基于在线可观测量的数据驱动关键尺度校准。在20个随机种子和85.6万次内层转移样本下,控制器使平均经验CVaR贝尔曼残差相比固定参数基线降低约85%(均值BEQ:1.2202降至0.1854;均值BEV:1.1624降至0.0535),并在不同CVaR水平、折扣因子和训练预算下维持稳定性。在时间顺序的样本外测试集上,所学策略实现夏普比率0.9281,最大回撤6.46%(含交易成本)。尽管买入持有策略累积回报更高(35.43% vs. 23.61%),但自适应策略显著更低波动率(9.57% vs. 47.93%)、回撤与CVaR损失。结果表明,仅通过改进训练过程而非改变风险目标,即可大幅提升风险敏感Q-learning在金融应用中的可靠性与风险调整后表现。

原文摘要 · Abstract (English)

Risk-aware Q-learning (RaQL) provides a model-free, two-timescale estimator for dynamic risk objectives, but its finite-budget behavior remains fragile: fixed inner-loop hyperparameters can produce unstable value estimates, persistent Bellman residuals, and inefficient sample reuse. This paper proposes an adaptive training controller for Conditional Value-at-Risk (CVaR) RaQL and evaluates it on a daily Bitcoin trading task. The controller preserves the original CVaR estimator and Bellman fixed point; instead, it redesigns the training procedure through six coordinated mechanisms: per-cell inner-step sizing, outer-rate-matched decay synchronization, a short early correction for the VaR-like inner variable, a coverage-first-then-greedy sample allocation rule, progressive suffix aggregation of mature inner estimates, and data-driven calibration of key scales from online-observable quantities. Across 20 random seeds and 856,000 inner-transition samples, the controller reduces the mean empirical CVaR Bellman residual by approximately 85% relative to the fixed-parameter baseline (MeanBEQ: 1.2202 to 0.1854; MeanBEV: 1.1624 to 0.0535) and maintains stability across CVaR levels, discount factors, and training budgets. On the chronological out-of-sample test set, the learned policy attains a Sharpe ratio of 0.9281 with a maximum drawdown of 6.46% after transaction costs. Although buy-and-hold yields a higher cumulative return (35.43% vs. 23.61%), the adaptive policy achieves far lower volatility (9.57% vs. 47.93%), drawdown, and CVaR loss. These results demonstrate that adaptive finite-budget training design, applied solely to the training procedure without altering the risk objective, can materially improve the reliability and risk-adjusted performance of risk-aware Q-learning in financial applications.

强化学习风险控制金融应用自适应训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。