用强化学习优化气候风险对冲,降低金融衍生品的碳成本
Climate-Dyna Deep Hedging for XVAs: Model-Based Reinforcement Learning, Residual Climate HVA, and Hedge-Instrument Discovery

- 基于配对世界模拟和动态修正,学习非线性对冲策略
- 碳成本从1.517降至0.831,接近理论最低0.821
- 仅用少量数据就能实现60.7%的精确收益,适合量化风控团队
对交易部门而言,残余气候对冲估值调整(HVA)是考虑现有对冲和可接受叠加策略后仍遗留的气候成本,无法通过单独压力损失推断。本文通过对比有/无气候影响的世界并重新优化叠加策略,计算该残余成本,并将对冲工具发现转化为估值问题:工具越能降低优化后的残余成本,就越有价值。在线性高斯情形下存在精确的有限时域Riccati解;在此基础上,Climate-Dyna从该对冲出发,通过成对世界模型滚动推演学习剩余非线性修正,并由独立门控决定是否部署更新。在公开数据校准的半合成欧盟碳交易体系研究中,计入原有对冲后,平均气候成本从1.517降至0.906,学习到的叠加策略进一步降至0.831,接近0.821的理论下限;残余Dyna相比回放方法仅需四分之一轨迹,就减少93%后悔值,且仅用25个目标转换即可保留60.7%的精确辅助收益。
原文摘要 · Abstract (English)
For a trading desk, residual climate hedging valuation adjustment (HVA) is the climate cost left after its inherited hedge and any admissible overlay have been taken into account; it therefore cannot be inferred from a stand-alone stress loss. We obtain this residual by comparing paired climate-on and baseline worlds and reoptimizing the overlay for each hedge universe, which also turns hedge-instrument discovery into a valuation problem: an instrument is useful to the extent that it lowers the optimized residual cost. The linear-Gaussian case has an exact finite-horizon Riccati solution; Climate-Dyna starts from that hedge and learns the remaining nonlinear correction from paired world-model rollouts, with an independent gate deciding whether to deploy the update. In a public-data-calibrated semi-synthetic EU ETS study, crediting the inherited hedge lowers the mean climate charge from 1.517 to 0.906, and the learned overlay lowers it to 0.831 against a 0.821 exact floor; residual Dyna cuts regret by 93% relative to replay with one quarter as many trajectories, while adaptation from only 25 target transitions retains 60.7% of the exact-assisted gain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。