arXiv:2601.03464cs.CL2026-01被引 1

LLM时间序列分类能力被提示评估严重低估,实际性能远超预期。

Prompting Underestimates LLM Capability for Time Series Classification

  • 用线性探针直接检测内部表征,发现模型已具备深层时序理解
  • 零样本提示F1仅0.15-0.26,线性探针提升至0.61-0.67
  • 早期变换层已存在类别判别性时序信息,适合研究模型内在表示

基于提示的评估显示大语言模型(LLMs)在时间序列分类任务上表现不佳,引发对其是否真正编码时间结构的质疑。我们证明这一结论源于提示生成方式的局限,而非模型表征能力不足:通过直接比较提示输出与同一内部表征的线性探针,发现零样本提示的平均F1仅为0.15-0.26,而线性探针将该值提升至0.61-0.67,常达到甚至超过专用时间序列模型性能。分层分析表明,类别判别性时序信息在早期变压器层中即已出现,并随视觉和多模态输入增强。结果揭示了当前评估方法与模型内在表示之间的系统性偏差,导致对模型时间序列理解能力的严重低估。

原文摘要 · Abstract (English)

Prompt-based evaluations suggest that large language models (LLMs) perform poorly on time series classification, raising doubts about whether they encode meaningful temporal structure. We show that this conclusion reflects limitations of prompt-based generation rather than the model's representational capacity by directly comparing prompt outputs with linear probes over the same internal representations. While zero-shot prompting performs near chance, linear probes improve average F1 from 0.15-0.26 to 0.61-0.67, often matching or exceeding specialized time series models. Layer-wise analyses further show that class-discriminative time series information emerges in early transformer layers and is amplified by visual and multimodal inputs. Together, these results demonstrate a systematic mismatch between what LLMs internally represent and what prompt-based evaluation reveals, leading current evaluations to underestimate their time series understanding.

大模型时间序列评估偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。