金融预测中,模型不确定时主动放弃判断,避免错误决策。
FinAbstain: Uncertainty-Calibrated Multimodal RAG for Selective Financial Forecasting
- 按时间点检索证据,分模块处理财报、新闻、技术指标等多源信息
- 仅当置信度达标才预测涨跌,否则拒答或转人工,降低错误率
- 首次系统整合多种不确定性校准方法,适合高风险金融场景
大型语言模型虽能合成金融叙事,但在证据不足、过时或矛盾时仍可能过度自信,这在预测中尤为危险。我们提出FinAbstain,一种不确定性校准的多模态检索增强生成框架,支持选择性预测。其时间点检索器仅采纳预测时刻公开的信息,并为基本面、新闻、技术、风险和验证代理提供模态专属证据。各代理的概率评估结合检索相关性、证据矛盾度、重复采样一致性及历史校准统计进行聚合。采用温度缩放、等向回归、容错预测及新提出的混合不确定性得分,在统一时间序列协议下评估。控制器仅在不确定性低于阈值时预测牛市、熊市或中性;否则主动放弃、请求证据、降低持仓或转交人工。评估涵盖一日与五日异常收益方向、二十日波动区间及弃权决策,使用准确率、校准性、风险-覆盖、引用、交易、延迟与成本等指标。为确保设计可审计,报告基于明确标注的模拟结果,而非实证主张。结果验证了核心假设:校准后的弃权策略可牺牲覆盖率以换取更低的定向误差与回撤。贡献包括时间安全架构、复合不确定性公式及可复现的评估范式,用于证据驱动的选择性金融预测。
原文摘要 · Abstract (English)
Large language models (LLMs) can synthesize financial narratives but may express high confidence when evidence is sparse, stale, or contradictory. This failure is especially consequential in forecasting, where filings, news, prices, volume, and technical signals can disagree. We present FinAbstain, a research framework for uncertainty-calibrated multimodal retrieval-augmented generation (RAG) with selective prediction. A point-in-time retriever admits only information public at the forecast timestamp and supplies modality-specific evidence to fundamental, news, technical, risk, and verification agents. Their probabilistic assessments are aggregated with retrieval relevance, evidence contradiction, repeated-sample consistency, and historical calibration statistics. Temperature scaling, isotonic regression, conformal prediction, and a proposed hybrid uncertainty score are evaluated under a common chronological protocol. A controller predicts bullish, bearish, or neutral outcomes only when uncertainty is below a validated threshold; otherwise it abstains, requests evidence, reduces exposure, or routes the case to human review. The evaluation covers one- and five-day abnormal-return direction, twenty-day volatility intervals, and abstention decisions, using accuracy, calibration, risk--coverage, citation, trading, latency, and cost metrics. To make the design auditable before a full data collection is complete, we report explicitly labeled simulated results rather than empirical claims. These results illustrate the intended hypothesis: calibrated abstention may trade coverage for lower selective error and drawdown. The contribution is a time-safe architecture, a composite uncertainty formulation, and a reproducible evaluation blueprint for evidence-grounded selective financial forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。