arXiv:2603.13252cs.AIcs.LG2026-03

提出双层不确定性机制,提升股票排名模型在市场突变时的安全性

When Alpha Breaks: Two-Level Uncertainty for Safe Deployment of Cross-Sectional Stock Rankers

  • 用排名位移预测构建事前安全基线,量化模型置信度
  • 20天策略下整体表现提升,尾部风险控制效果显著
  • 适合量化交易系统中需防范信号失效的场景

横截面股票排名模型常被当作点预测使用:模型输出得分,投资组合按排序执行。但在非平稳环境下,模型在制度转换时可能失效。在AI股票预测器中,一个LightGBM排名模型在20天预测周期上总体表现良好,但2024年回测期恰逢人工智能主题上涨和板块轮动,导致长期信号失效且20天信号减弱。这促使我们将部署拆分为两个决策:(i) 策略是否应交易;(ii) 如何控制活跃交易中的风险。我们改进了直接认知不确定性预测(DEUP),通过预测排名位移并定义相对于时点基线(PIT-safe)的认知不确定性信号 ehat。实证显示,ehat与信号强度结构相关(1,865个日期中,ehat与绝对得分的中位相关性约为0.6),因此反向不确定性加权会削弱强信号并降低表现。为此,我们提出两层部署策略:策略级的制度信任门 G(t),决定是否交易(总体AUROC约0.72,最终阶段为0.75);以及位置级的认知尾部风险上限,仅对最不确定预测降仓。操作策略为:当G(t) ≥ 0.2时才交易,活跃日进行波动率缩放,并对最高认知尾部进行封顶。该策略在20天对比中提升了风险调整后收益,表明DEUP主要作为尾部风险防护而非连续加权因子发挥作用。

原文摘要 · Abstract (English)

Cross-sectional ranking models are often deployed as if point predictions were sufficient: the model outputs scores and the portfolio follows the induced ordering. Under non-stationarity, rankers can fail during regime shifts. In the AI Stock Forecaster, a LightGBM ranker performs well overall at a 20-day horizon, yet the 2024 holdout coincides with an AI thematic rally and sector rotation that breaks the signal at longer horizons and weakens 20d. This motivates treating deployment as two decisions: (i) whether the strategy should trade at all, and (ii) how to control risk within active trades. We adapt Direct Epistemic Uncertainty Prediction (DEUP) to ranking by predicting rank displacement and defining an epistemic uncertainty signal ehat relative to a point-in-time (PIT-safe) baseline. Empirically, ehat is structurally coupled with signal strength (median correlation between ehat and absolute score is about 0.6 across 1,865 dates), so inverse-uncertainty sizing de-levers the strongest signals and degrades performance. To address this, we propose a two-level deployment policy: a strategy-level regime-trust gate G(t) that decides whether to trade (AUROC around 0.72 overall and 0.75 in FINAL) and a position-level epistemic tail-risk cap that reduces exposure only for the most uncertain predictions. The operational policy, trade only when G(t) is at least 0.2, apply volatility sizing on active dates, and cap the top epistemic tail, improves risk-adjusted performance in the 20d policy comparison and indicates DEUP adds value mainly as a tail-risk guard rather than a continuous sizing denominator.

量化交易不确定性风险管理排名模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。