通过匹配任务频谱特性,让状态空间模型更高效学习少量数据。
Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
- 基于频谱匹配设计任务相关初始化方法
- 在频谱不匹配时显著提升小样本泛化性能
- 适合追求数据效率的序列建模研究者
现代人工智能的成功与规模定律密切相关,但高质量数据有限,数据效率——用更少数据学得更多——成为关键挑战。模型的归纳偏置是提升数据效率的重要杠杆,但经典序列模型如状态空间模型(SSMs)通常依赖固定、任务无关的先验。当这种固定偏置与任务真实结构不匹配时,模型需额外样本才能克服自身偏差。本文提出一种理论框架,用于理解并对齐线性时不变SSM的归纳偏置。通过分析SSM诱导的核函数,发现其谱特性由模型频率响应决定。据此提出任务依赖初始化(TDI),一种快速的功率谱匹配方法,在下游训练前将SSM初始偏置与任务谱特征对齐。在受控合成实验、可训练单层SSM及多种真实世界基准的深层SSM上,当任务存在相关谱结构且默认偏置谱不匹配时,TDI显著提升数据高效泛化能力。结果为任务自适应归纳偏置提供了理论视角与实用工具,指明了更高效序列建模的新路径。
原文摘要 · Abstract (English)
The remarkable success of modern AI has been closely tied to scaling laws, yet the finite supply of high-quality data makes data efficiency--learning more from less--an increasingly important frontier. A model's inductive bias is a critical lever for data efficiency, but foundational sequence models such as State Space Models (SSMs) often rely on fixed, task-agnostic biases. When this fixed prior is misaligned with the underlying structure of a task, the model may require additional samples to overcome its own bias before learning the relevant signal. In this work, we introduce a principled framework for understanding and aligning the inductive bias of linear time-invariant SSMs. We first formalize this bias through an SSM-induced kernel and show theoretically and empirically that its spectrum is governed by the model's frequency response. This characterization motivates Task-Dependent Initialization (TDI), a fast power-spectrum matching method that aligns the initial SSM bias with the task's spectral characteristics before downstream training. Across controlled synthetic experiments, trainable one-layer SSMs, and deep SSMs on diverse real-world benchmarks, TDI can improve data-efficient generalization primarily when task-relevant spectral structure is present and the default SSM bias is spectrally mismatched. Our results provide both a theoretical lens and a practical tool for task-adaptive inductive bias, suggesting a path toward more data-efficient sequence modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。