用不确定性感知框架提升高风险决策的稳定性与收益
UAMDP: Uncertainty-Aware Markov Decision Process for Risk-Constrained Reinforcement Learning from Probabilistic Forecasts
- 融合贝叶斯预测与后验采样,动态管理未知风险
- 交易策略夏普比率升至1.74,最大回撤减半
- 适合金融、供应链等高波动场景的智能决策
在高风险、高波动的序列决策场景中,单纯追求期望回报不足以为据;必须系统性管理不确定性。本文提出不确定性感知马尔可夫决策过程(UAMDP),统一整合贝叶斯预测、后验采样强化学习与条件风险价值(CVaR)约束下的规划。闭环中,智能体更新对潜在动态的认知,通过Thompson采样生成可能未来路径,并在预设风险容忍度下优化策略。我们建立了收敛至贝叶斯最优基准的后悔界。在高频股票交易与零售库存控制两个领域评估,相较于强基线深度学习方法,UAMDP将长期预测均方根误差降低最多25%,对称平均绝对百分比误差降低32%;交易夏普比率从1.54提升至1.74,最大回撤约减半。结果表明,结合校准的概率建模、与后验不确定性对齐的探索以及风险敏感控制,可实现更安全、更盈利的通用序列决策。
原文摘要 · Abstract (English)
Sequential decisions in volatile, high-stakes settings require more than maximizing expected return; they require principled uncertainty management. This paper presents the Uncertainty-Aware Markov Decision Process (UAMDP), a unified framework that couples Bayesian forecasting, posterior-sampling reinforcement learning, and planning under a conditional value-at-risk (CVaR) constraint. In a closed loop, the agent updates its beliefs over latent dynamics, samples plausible futures via Thompson sampling, and optimizes policies subject to preset risk tolerances. We establish regret bounds that converge to the Bayes-optimal benchmark under standard regularity conditions. We evaluate UAMDP in two domains including high-frequency equity trading and retail inventory control, both marked by structural uncertainty and economic volatility. Relative to strong deep learning baselines, UAMDP improves long-horizon forecasting accuracy (RMSE decreases by up to 25% and sMAPE by 32%), and these gains translate into economic performance: the trading Sharpe ratio rises from 1.54 to 1.74 while maximum drawdown is roughly halved. These results show that integrating calibrated probabilistic modeling, exploration aligned with posterior uncertainty, and risk-aware control yields a robust, generalizable approach to safer and more profitable sequential decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。