让量化交易策略在风险敏感下保持安全,通过可解释的约束优化实现零违规。
Tail-Safe Hedging: Explainable Risk-Sensitive Reinforcement Learning with a White-Box CBF--QP Safety Layer in Arbitrage-Free Markets
- 用分布式强化学习结合尾部风险控制,动态调整极端情况下的策略采样。
- 在模拟市场中降低左尾风险,且只要求解可行即实现零硬约束违规。
- 提供完整运行日志供审计,适合金融风控与合规监管场景使用。
我们提出 Tail-Safe,一种面向部署的衍生品对冲框架,融合分布式、风险敏感强化学习与针对金融约束定制的白盒控制屏障函数-二次规划(CBF-QP)安全层。学习部分采用基于IQN的分布评论器与CVaR目标(IQN--CVaR--PPO),并引入尾部覆盖率控制器,通过温度倾斜和尾部增强调节分位数采样,稳定小α值下的估计。安全层强制执行离散时间CBF不等式及领域特定约束——椭球无交易带、箱型与速率限制、符号一致性门控,以凸二次规划求解,其遥测数据(活跃集、紧度、速率利用率、门控评分、松弛量与求解状态)构成可审计的治理轨迹。理论贡献包括:在模型偏差有界下保证安全集的鲁棒前向不变性;QP的最小偏离投影解释;从每状态KL正则化到最坏情况CVaR的KL-to-DRO上界;温度倾斜CVaR估计器的集中性与样本复杂度结果;在KL约束下基于CVaR的信任区域改进不等式;以及到期感知收紧下的可行性持续性。实验表明,在无套利、微观结构感知的合成市场(SSVI→Dupire→VIX,ABIDES/MockLOB执行)中,Tail-Safe在不损害中心性能的前提下有效降低左尾风险,且当QP可行且松弛为零时,始终实现零硬约束违规。遥测数据映射至治理仪表盘与事件工作流,支持可解释性与可审计性。局限在于依赖合成数据与简化执行以隔离方法贡献。
原文摘要 · Abstract (English)
We introduce Tail-Safe, a deployability-oriented framework for derivatives hedging that unifies distributional, risk-sensitive reinforcement learning with a white-box control-barrier-function (CBF) quadratic-program (QP) safety layer tailored to financial constraints. The learning component combines an IQN-based distributional critic with a CVaR objective (IQN--CVaR--PPO) and a Tail-Coverage Controller that regulates quantile sampling through temperature tilting and tail boosting to stabilize small-$α$ estimation. The safety component enforces discrete-time CBF inequalities together with domain-specific constraints -- ellipsoidal no-trade bands, box and rate limits, and a sign-consistency gate -- solved as a convex QP whose telemetry (active sets, tightness, rate utilization, gate scores, slack, and solver status) forms an auditable trail for governance. We provide guarantees of robust forward invariance of the safe set under bounded model mismatch, a minimal-deviation projection interpretation of the QP, a KL-to-DRO upper bound linking per-state KL regularization to worst-case CVaR, concentration and sample-complexity results for the temperature-tilted CVaR estimator, and a CVaR trust-region improvement inequality under KL limits, together with feasibility persistence under expiry-aware tightening. Empirically, in arbitrage-free, microstructure-aware synthetic markets (SSVI $\to$ Dupire $\to$ VIX with ABIDES/MockLOB execution), Tail-Safe improves left-tail risk without degrading central performance and yields zero hard-constraint violations whenever the QP is feasible with zero slack. Telemetry is mapped to governance dashboards and incident workflows to support explainability and auditability. Limitations include reliance on synthetic data and simplified execution to isolate methodological contributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。