arXiv:2512.00366cs.LGcs.AI2025-12AAAI被引 5

用语义与频谱联合知识蒸馏,让轻量模型预测更准更合理。

S^2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting

  • 融合视觉频谱与文本语义,构建双重视角的知识蒸馏框架
  • 在WeatherBench和TaxiBJ+上使轻量学生模型超越先进方法
  • 适合需要高精度长期预测的气象、交通等场景

时空预测常依赖计算密集型模型以捕捉复杂动态。知识蒸馏(KD)已成为构建轻量级学生模型的关键技术,近期如频域感知KD的方法能有效保留频谱特性(如高频细节和低频趋势)。然而,这些方法仅在像素层面操作,无法感知视觉模式背后的丰富语义与因果关系。为此,我们提出S^2-KD,一种将语义先验与频谱表示统一用于蒸馏的新框架。该方法首先训练一个具备多模态能力的教师模型,利用大语言多模态模型(LMM)提供的文本叙事来推理事件成因,同时其架构在潜在空间中解耦频谱成分。核心是新的蒸馏目标,将这种统一的语义-频谱知识迁移到仅视觉输入的轻量学生模型中。结果是,学生模型的预测不仅频谱准确,且语义连贯,推理时无需文本输入或额外结构开销。在WeatherBench和TaxiBJ+等基准上的大量实验表明,S^2-KD显著提升简单学生模型性能,使其在长时程和复杂非平稳场景下超越现有先进方法。

原文摘要 · Abstract (English)

Spatiotemporal forecasting often relies on computationally intensive models to capture complex dynamics. Knowledge distillation (KD) has emerged as a key technique for creating lightweight student models, with recent advances like frequency-aware KD successfully preserving spectral properties (i.e., high-frequency details and low-frequency trends). However, these methods are fundamentally constrained by operating on pixel-level signals, leaving them blind to the rich semantic and causal context behind the visual patterns. To overcome this limitation, we introduce S^2-KD, a novel framework that unifies Semantic priors with Spectral representations for distillation. Our approach begins by training a privileged, multimodal teacher model. This teacher leverages textual narratives from a Large Multimodal Model (LMM) to reason about the underlying causes of events, while its architecture simultaneously decouples spectral components in its latent space. The core of our framework is a new distillation objective that transfers this unified semantic-spectral knowledge into a lightweight, vision-only student. Consequently, the student learns to make predictions that are not only spectrally accurate but also semantically coherent, without requiring any textual input or architectural overhead at inference. Extensive experiments on benchmarks like WeatherBench and TaxiBJ+ show that S^2-KD significantly boosts the performance of simple student models, enabling them to outperform state-of-the-art methods, particularly in long-horizon and complex non-stationary scenarios.

知识蒸馏时空预测多模态频谱建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。