测试金融预测中模型的可信度与不确定性,避免过度自信导致亏损。
FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting
- 设计时间严格约束的评估框架,防止信息泄露。
- 用概率评分和区间得分衡量模型信心是否匹配实际表现。
- 适合关注模型可靠性与风险控制的量化研究者。
大型语言模型正越来越多地被用于具备观察、规划与执行能力的智能体系统中。在金融领域,即使只是辅助性的输出,一旦被用来决定交易规模或风险分配,也会直接影响决策结果。一个关键问题在于信心与能力之间的差距:模型表现仅略高于随机水平,却持续高估自身能力,按常规投注规则将导致长期收益为负。现有基准侧重语义理解或点预测精度,但未直接测试在真实市场的时间约束与非平稳性条件下的概率校准能力。本文提出FinBench,一个专为金融预测设计的校准与不确定性评估基准,其特点为(i)严格时间门控以避免前瞻偏差,(ii)采用严格合理的评分规则惩罚虚构的信心。任务要求模型输出(a)正收益的概率,(b)实际对数收益率的80%预测区间;评估使用Brier评分与Winkler区间评分,并对比硬性基线的技能分数。本文描述了基准规范,并报告了一次小规模试点运行(1个交易日;3只流动性标的;33次预测),作为流程验证。试点结果表明,校准敏感指标可有效区分‘自信但脆弱’行为与具备不确定性意识的预测模式。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act. In finance, even "assistive" systems become decision-relevant once their outputs are used to size trades or allocate risk. A key failure mode is the confidence--competence gap: a model that is only slightly better than chance but consistently overconfident will, under typical bet-sizing rules, generate negative long-run growth. Existing benchmarks emphasize semantic understanding or point accuracy, but do not directly test probabilistic calibration under the temporal constraints and non-stationarity that define real markets. We introduce FinBench, a benchmark designed to evaluate calibration and uncertainty quality for financial forecasting in a setting that is (i) strictly time-gated to avoid look-ahead bias and (ii) evaluated with strictly proper scoring rules that penalize hallucinated confidence. FinBench tasks require models to output (a) a probability of positive return and (b) an 80% prediction interval for realized log return; evaluation uses the Brier score and the Winkler interval score, along with skill scores against hard baselines. This paper describes the benchmark specification and reports a small pilot run (one trading day; three liquid tickers; 33 forecasts) as a sanity check of the pipeline. The pilot illustrates how calibration-sensitive metrics distinguish between "confident but fragile" behavior and uncertainty-aware forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。