金融新闻预测模型的性能夸大,因时间泄漏导致结果不可靠。
Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal
- 用时序划分数据集,发现模型性能被严重高估。
- 并购新闻在时序验证下仍具显著预测信号(MCC=0.068)。
- 建议将泄漏审计作为金融NLP基准的必备披露。
金融新闻方向预测已成为热门NLP基准,但其表现高度依赖训练测试集是否按时间顺序划分,即是否存在时间泄漏。我们对49,799篇新闻文章、16种特征-模型组合(涵盖TF-IDF、MiniLM、FinBERT、微调的RoBERTa-large/DeBERTa-v3-large,以及Llama-3/Qwen2.5的零样本/少样本和LoRA探针)进行了多架构审计。随机划分使马修相关系数(MCC)虚增1.1倍至6.5倍,且与模型容量和特征丰富度正相关;端到端微调的FinBERT反而进一步放大差距(比例达1.75倍)。按事件类型分析,只有并购(M&A)在近时序划分下仍具正向锁定测试信号(TF-IDF MCC=0.138训练仅用,0.068训练+验证集重拟合;10,000次置换检验p<10⁻³),该信号未迁移至FNSPID的2009–2020年美国语料,表明其局限于2024–2025年欧洲导向的并购语义,非普适预测器。三个独立角色标注器一致指向收购方相关文章为信号源,属功率受限的定性收敛,非假设检验的不对称性。时序划分在金融NLP中的作用,如同资产定价中的特征清洗:剔除可预测的陈旧信息,留下小规模、事件局部化、词汇浅层的残差。我们主张将泄漏审计作为金融NLP基准的强制披露要求。
原文摘要 · Abstract (English)
Financial-news direction prediction has become a popular NLP benchmark, yet reported gains depend critically on whether the train-test split is chronological or random, i.e., on temporal leakage. We audit this dependence on a 49,799-article corpus across 16 feature-model combinations spanning TF-IDF, MiniLM, FinBERT, and fine-tuned RoBERTa-large / DeBERTa-v3-large, plus separate zero/few-shot and LoRA probes of Llama-3 and Qwen2.5 LLMs: random splits inflate MCC by $1.1\times$ to $6.5\times$, tracking model capacity and feature richness, and end-to-end FinBERT fine-tuning re-amplifies rather than closes the gap (size-matched ratio $1.75\times$). Conditioning on event type, mergers and acquisitions (M&A) is the only audited category with a positive locked-test signal under near-temporal chronological evaluation (TF-IDF MCC $= 0.138$ train-only, $0.068$ under train$\cup$val refit; 10,000-permutation $p < 10^{-3}$); the signal does not transfer to FNSPID's 2009-2020 U.S. corpus, localising the headline to our 2024-2025 European-tilted M&A semantics rather than a universal predictor. Three independent role labellers converge on acquirer-tagged articles as the signal locus, a power-limited qualitative convergence rather than a hypothesis-tested asymmetry. Chronological splitting plays for financial NLP the role characteristics-purging plays for asset pricing: it strips the predictable, stale component of news and leaves a residual that is small, event-localized, and lexically shallow. We advocate leakage audits as a required disclosure for financial-NLP benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。