用可微分的复合奖励函数平衡收益与风险,支持个性化投资偏好。
A Risk-Aware Reinforcement Learning Reward for Financial Trading
- 设计四部分可微奖励:年化收益、下行风险、差额收益和特雷诺比率
- 通过权重调节实现不同风险收益目标,网格搜索优化参数
- 支持扩展新风险度量或自适应权重,适合复杂交易场景
我们提出一种用于金融交易强化学习的新型复合奖励函数,通过四个可微项平衡收益与风险:年化收益、下行风险、差额收益和特雷诺比率。与单一指标(如夏普比率)相比,该框架模块化且由权重w1、w2、w3、w4参数化,可编码多样投资者偏好。通过网格搜索调整权重以实现特定风险-收益特征。为每项推导闭式梯度,支持基于梯度的训练,并分析了单调性、有界性和模块性等关键理论性质。该框架为复杂交易环境中构建稳健多目标奖励函数提供了通用范式,可扩展至其他风险度量或自适应权重机制。
原文摘要 · Abstract (English)
We propose a novel composite reward function for reinforcement learning in financial trading that balances return and risk using four differentiable terms: annualized return downside risk differential return and the Treynor ratio Unlike single metric objectives for example the Sharpe ratio our formulation is modular and parameterized by weights w1 w2 w3 and w4 enabling practitioners to encode diverse investor preferences We tune these weights via grid search to target specific risk return profiles We derive closed form gradients for each term to facilitate gradient based training and analyze key theoretical properties including monotonicity boundedness and modularity This framework offers a general blueprint for building robust multi objective reward functions in complex trading environments and can be extended with additional risk measures or adaptive weighting
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。