arXiv:2607.14373cs.LGq-fin.RM2026-07

用逆强化学习从噪声决策中提取风险偏好并优化策略。

A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning

  • 基于自适应贝叶斯逆强化学习,从含噪行为中推断风险偏好。
  • 算法收敛率达 $O(\exp(-cm+O(\sqrt{m\log m})))$,理论保证强。
  • 适用于金融等复杂场景,支持多种风险度量的统一优化。

我们提出一种抗噪声的“诱使-优化”框架,结合逆强化学习(IRL)与强化学习(RL),用于在广义扭曲风险度量下获取智能体的风险偏好并优化策略。在偏好获取方面,提出一种自适应贝叶斯IRL方法,从智能体的含噪观测决策中推断其潜在风险目标,明确允许随机及次优动作。我们证明,在候选类中存在有限个区分性问题可识别最优扭曲风险度量,并在一般设定下建立了算法收敛率为 $O(\exp(-cm+O(\sqrt{m\log m})))$,其中 $c>0$ 为常数,$m$ 为迭代次数。在优化方面,开发了无模型强化学习算法以实现条件扭曲风险度量下的策略优化。通过将目标表示为条件代价分位函数对扭曲函数的积分,该方法统一了各类扭曲风险目标。通过扩展近端策略优化(PPO)算法,引入策略、价值与分位数神经网络,其中分位数网络估计完整条件代价分位函数,支持一般风险目标的数值评估。实证研究全面验证了该框架在复杂金融环境中的偏好获取准确性和优化有效性。

原文摘要 · Abstract (English)

We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents' risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics. On the elicitation side, we propose an adaptive Bayesian IRL method that infers agents' latent risk objectives from their noisy observed decisions, explicitly allowing agents to take stochastic and suboptimal actions. We establish the existence of a finite set of distinguishing questions that identifies the preferred distortion riskmetric within the candidate class and prove that the convergence rate of the algorithm is of order $O(\exp(-cm+O(\sqrt{m\log m})))$ under general settings, where $c>0$ is a constant and $m$ denotes the number of algorithm iterations. On the optimization side, we develop a model-free RL algorithm for optimizing policies under conditional distortion riskmetrics. By representing the objective as an integral of the conditional cost quantile function with respect to the distortion function, the method unifies distortion-riskmetric objectives. We optimize diverse risk objectives by extending the Proximal Policy Optimization (PPO) algorithm with policy, value, and quantile neural networks, where the quantile network estimates the full conditional cost quantile function and enables numerical evaluation of general risk objectives. A comprehensive empirical study demonstrates the framework's elicitation accuracy and effectiveness in complex financial environments.

强化学习风险建模逆强化学习金融优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。