arXiv:2509.25696cs.LGcs.CL2025-09

用视觉语言模型生成伪标签,训练出比原模型更强的时间序列问答模型。

Can VLM Pseudo-Labels Train a Time-Series QA Model That Outperforms the VLM?

  • 利用视觉语言模型生成时间序列的伪标签进行训练。
  • 模型在大量无标注数据上训练后,性能超越原始视觉语言模型。
  • 适合缺乏标注数据的时间序列问答任务研究者。

时间序列问答(TSQA)任务因标注数据稀缺而面临挑战。近年来,大规模视觉语言模型(VLM)展现出零样本分析时间序列信号的潜力。本文提出一种利用VLM生成伪标签的训练方法。尽管VLM可能产生错误标签,但得益于深度神经网络对噪声标签的内在鲁棒性,基于这些伪标签仍可有效训练TSQA模型。实验结果表明,使用大量无标注数据的模型不仅成功训练,且性能超越了原始的VLM。该方法在数据有限场景下显著提升了问答能力。

原文摘要 · Abstract (English)

Time-series question answering (TSQA) tasks face significant challenges due to the lack of labeled data. Alternatively, with recent advancements in large-scale models, vision-language models (VLMs) have demonstrated the potential to analyze time-series signals in a zero-shot manner. In this paper, we propose a training approach that uses pseudo labels generated by a VLM. Although VLMs can produce incorrect labels, TSQA models can still be effectively trained based on the property that deep neural networks are inherently robust to such noisy labels. Our experimental results demonstrate that TSQA models are not only successfully trained with pseudo labels, but also surpass the performance of the VLM itself by leveraging a large amount of unlabeled data.

时间序列伪标签VLM问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。