arXiv:2510.04563cs.LGmath.OC2025-10

提出优化风险度量的新算法,提升投资组合鲁棒性。

Stochastic Approximation Methods for Distortion Risk Measure Optimization

  • 用双重表示法设计梯度下降算法,分时尺度跟踪量化值并更新决策变量。
  • DM形式收敛率达O(k⁻⁴ᐟ⁷),QF形式更快达O(k⁻²ᐟ³)。
  • 适用于强化学习中的库存管理,兼顾稳定性与计算效率。

扭曲风险度量(DRMs)能捕捉决策中的风险偏好,是管理不确定性的通用准则。本文基于两种对偶表示——扭曲度量(DM)形式和分位数函数(QF)形式——提出梯度下降优化算法。DM形式采用三时尺度算法,结合广义似然比与核密度估计,追踪分位数、计算梯度并更新决策变量;QF形式则采用更简单的双时尺度方法,避免复杂分位数梯度估计。混合形式结合两者优势,在扭曲函数突变处使用DM形式保证稳健性,平滑区域使用QF形式提高效率。理论证明了算法的强收敛性及收敛速率:DM形式达到最优率O(k⁻⁴ᐟ⁷),QF形式更快为O(k⁻²ᐟ³)。数值实验验证其有效性,在鲁棒投资组合选择任务中显著优于基线方法。方法可扩展至深度强化学习,本文开发了基于DRM的近端策略优化算法,并应用于多级动态库存管理,展现实际应用价值。

原文摘要 · Abstract (English)

Distortion Risk Measures (DRMs) capture risk preferences in decision-making and serve as general criteria for managing uncertainty. This paper proposes gradient descent algorithms for DRM optimization based on two dual representations: the Distortion-Measure (DM) form and Quantile-Function (QF) form. The DM-form employs a three-timescale algorithm to track quantiles, compute their gradients, and update decision variables, utilizing the Generalized Likelihood Ratio and kernel-based density estimation. The QF-form provides a simpler two-timescale approach that avoids the need for complex quantile gradient estimation. A hybrid form integrates both approaches, applying the DM-form for robust performance around distortion function jumps and the QF-form for efficiency in smooth regions. Proofs of strong convergence and convergence rates for the proposed algorithms are provided. In particular, the DM-form achieves an optimal rate of $O(k^{-4/7})$, while the QF-form attains a faster rate of $O(k^{-2/3})$. Numerical experiments confirm their effectiveness and demonstrate substantial improvements over baselines in robust portfolio selection tasks. The method's scalability is further illustrated through integration into deep reinforcement learning. Specifically, a DRM-based Proximal Policy Optimization algorithm is developed and applied to multi-echelon dynamic inventory management, showcasing its practical applicability.

风险优化强化学习投资组合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。