arXiv:2603.18688cs.LGcs.CL2026-03

用跨领域知识蒸馏,让科学时序数据学会通用表示。

STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation

  • 通过跨域知识蒸馏融合多类时序模型优势
  • 在7个科学任务上显著提升表示能力
  • 适合需要通用时序表征的科研场景

科学时序数据是科学AI的核心,但普遍存在稀疏、异构和规模小的问题,统一表征学习极具挑战。与此同时,音频、通用时序和脑信号等领域的基础模型蕴含丰富知识,但其在科学信号上的适用性尚未充分探索。本文系统评估了相关基础模型,发现其知识迁移有效且具有互补性。基于此,提出STEP框架——通过跨域蒸馏构建科学时序编码器。STEP引入自适应分块处理极端长序列,采用统计补偿机制应对不同数值尺度,并融合多领域知识,学习适配科学信号的通用可迁移特征。在7个科学时序任务上的实验表明,STEP不仅提供有效结构,更建立有效预训练范式,推动科学时序表示学习向前迈出关键一步。

原文摘要 · Abstract (English)

Scientific time series are central to scientific AI but are typically sparse, highly heterogeneous, and limited in scale, making unified representation learning particularly challenging. Meanwhile, foundation models pretrained on relevant time series domains such as audio, general time series, and brain signals contain rich knowledge, but their applicability to scientific signals remains underexplored. In this paper, we investigate the transferability and complementarity of foundation models from relevant time series domains, and study how to effectively leverage them to build a unified encoder for scientific time series. We first systematically evaluate relevant foundation models, showing the effectiveness of knowledge transfer to scientific tasks and their complementary strengths. Based on this observation, we propose STEP, a Scientific Time Series Encoder Pretraining framework via cross domain distillation. STEP introduces adaptive patching to handle extreme-length sequences and a statistics compensation scheme to accommodate diverse numerical scales. It further leverages cross-domain distillation to integrate knowledge from multiple foundation models into a unified encoder. By combining complementary representations across different domains, STEP learns general-purpose and transferable features tailored for scientific signals. Experiments on seven scientific time series tasks demonstrate that STEP provides both an effective structure and an effective pretraining paradigm, taking a STEP toward scientific time series representation learning.

时序建模知识蒸馏科学计算预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。