arXiv:2607.12248q-fin.STcs.LG2026-07

揭露股价预测中方向准确率的假象,提出防陷阱的基准测试方法

When Directional Accuracy Lies: A Base-Rate-Honest Benchmark for LoRA-Adapted TimesFM on Equity Forecasting

论文配图:When Directional Accuracy Lies: A Base-Rate-Honest Benchmark for LoRA-Adapted TimesFM on Equity Forecasting
图 1 · 摘自论文原文
  • 构建滚动验证、分层留出的严谨基准,杜绝市场上升带来的误判
  • 80%准确率在牛市中几乎纯靠‘永远上涨’规则就能达成,模型实际无真本领
  • 多组实验反复验证:微调模型无方向性优势,仅降低点预测误差

大型预训练时序模型如TimesFM在金融预测中备受关注,但原始方向准确率在股市中具有误导性。本研究发现,早期采用LoRA适配器的模型看似达到约80%方向准确率,实则并非能力体现。在长期牛市中,一个不使用输入的‘始终上涨’规则也能获得相近准确率。为区分真实技能与基础率偏差,我们构建了一个可复现、数据冻结的基准:采用扩展式走查折叠、分层留出股票划分、诚实基线(零样本TimesFM、始终上涨、随机游走、持久性、AR(1))及受Benjamini-Hochberg FDR控制的成对显著性检验(McNemar、Diebold-Mariano)。在科技股集中的纳斯达克100和广义标普500两个标的中均应用该方法,报告相对于‘始终上涨’基线的超额准确率。三项发现重复成立:一、重现历史约80%条件后,真实基础率约为0.70,而微调模型得分低于此值;二、无论哪个标的,在任意时间跨度上,合并的LoRA适配器均无方向性优势(六个月期为负);三、按行业专业化适配效果显著劣于单一合并适配器(在持有期股票上,Diebold-Mariano p<0.001)。微调唯一可观测收益是点预测误差的统计学降低,但仍不及朴素基线,且未带来可交易的方向性优势。贡献在于方法论:提供一套防基础率陷阱的可信赖协议,以及其产出的可复现负面结果。

原文摘要 · Abstract (English)

Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard in equity markets. An early LoRA adapter in this project appeared to reach roughly 80% directional accuracy; we show this is not evidence of skill. Over a long horizon in a rising market, a trivial "always-up" rule attains comparably high accuracy without using the input at all. To separate genuine skill from this base-rate artifact, we build a reproducible, frozen-data benchmark with expanding walk-forward folds, a stratified held-out-ticker split, honest baselines (zero-shot TimesFM, always-up, random-walk, persistence, AR(1)), and paired significance tests (McNemar, Diebold-Mariano) under Benjamini-Hochberg FDR control. We apply the identical method to two universes -- a tech-heavy NASDAQ-100 and a broad S&P 500 -- reporting excess accuracy over the always-up base rate. Three findings replicate. First, when the historical ~80% condition is recreated, the high number is a base rate of ~0.70 that the fine-tuned model scores below. Second, pooled LoRA shows no directional skill over the base rate at any horizon on either universe (negative at the six-month horizon). Third, per-sector specialization is significantly worse than a single pooled adapter (Diebold-Mariano p<0.001 on held-out stocks at h=128). Fine-tuning's only measurable benefit is a statistically significant reduction in point-forecast error relative to zero-shot TimesFM, which nonetheless does not beat naive baselines and confers no tradeable directional edge. The contribution is methodological: a defensible, fully seeded protocol that prevents the base-rate trap, together with the replicated negative result it produces.

金融预测基准测试LoRA适配方向准确率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。