arXiv:2608.22321cs.CL2026-08

测试发现文本内容对多模态时序预测效果影响极小,实际提升来自数值特征。

Semantics or Structure? Auditing Text Sensitivity in Multimodal Time-Series Forecasting

  • 通过替换文本内容验证模型是否真依赖语义信息。
  • 更换任意文本后误差变化不足0.5%,说明文本非关键信号。
  • 适合研究多模态模型真实依赖机制的学者参考。

多模态时序预测中,自然语言上下文被认为能提升预测性能。近期模型如Aurora、MM-TSFlib和TaTS在Time-MMD基准上相比单模态基线取得显著提升,归因于文本信息。然而,这些模型是否真正敏感于文本语义仍未经验证。本文通过控制性文本扰动、归因分析及对Aurora文本路径的探测,发现将Time-MMD中每行文本替换为任意真实文本(空白、恒定、同域乱序或跨域)后,三种架构的均方误差变化均小于0.5%。当移除共提供数值列而未修改文本时,文献中报告的性能提升依然恢复。结论表明,在此基准及此类冻结编码器架构下,文本内容并非报告性能提升的实际信号。为支持未来在结构化数据多模态基础模型中融合文本的研究,我们发布可复用的扰动协议与评估工具包作为诊断工具。

原文摘要 · Abstract (English)

Multimodal time-series forecasting has emerged as a promising paradigm in which natural-language context is expected to improve predictive performance. Recent multimodal foundation models, including Aurora, as well as early- and late-fusion approaches such as MM-TSFlib and TaTS, report substantial gains over unimodal baselines on the Time-MMD benchmark, attributing these improvements to textual information. However, whether these models are actually sensitive to the semantic content of the text remains unverified. We address this question through controlled text perturbations, attribution analyses, and probes of Aurora's text pathway. On Time-MMD, swapping each row's text for any other real text (empty, constant, within-domain shuffled, or cross-domain) moves mean MSE by less than $0.5\%$ on all three architectures. The improvement reported in the literature is recovered when a co-shipped numeric column is removed without touching text. We conclude that, on this benchmark and within this family of frozen-encoder architectures, text content is not the operative signal behind the reported gains. To support future work on text integration in multimodal foundation models for structured data, we release our perturbation protocol and evaluation harness as a reusable diagnostic toolkit.

时序预测多模态文本敏感性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。