arXiv:2602.13550cs.LG2026-02被引 1

让神经网络在训练外数据上更可靠,通过权重空间序列建模提升泛化能力。

Out-of-Support Generalisation via Weight-Space Sequence Modelling

  • 将模型权重变化视为序列,按训练数据分布分层建模
  • 在合成数据和空气质量数据上表现优于或媲美现有方法
  • 无需额外假设即可生成可信且带不确定性的预测,适合安全关键场景

随着深度学习在关键领域的突破,模型需对训练集范围外的数据进行外推,这一挑战被称为超出支持范围(OoS)泛化。然而神经网络在OoS样本上常出现灾难性失效,产生看似合理却过度自信的预测。本文将OoS泛化重构为权重空间中的序列建模问题,将训练集划分为对应离散步骤的同心壳层。WeightCaster框架在不依赖显式归纳偏置的前提下,实现可解释、可信且带有不确定性感知的预测,同时保持高计算效率。在合成余弦数据集和真实空气质量传感器数据上的实证验证表明,其性能达到或超过当前最优水平。该方法显著提升了模型在分布外场景下的可靠性,对人工智能在安全关键应用中的广泛采用具有重要意义。

原文摘要 · Abstract (English)

As breakthroughs in deep learning transform key industries, models are increasingly required to extrapolate on datapoints found outside the range of the training set, a challenge we coin as out-of-support (OoS) generalisation. However, neural networks frequently exhibit catastrophic failure on OoS samples, yielding unrealistic but overconfident predictions. We address this challenge by reformulating the OoS generalisation problem as a sequence modelling task in the weight space, wherein the training set is partitioned into concentric shells corresponding to discrete sequential steps. Our WeightCaster framework yields plausible, interpretable, and uncertainty-aware predictions without necessitating explicit inductive biases, all the while maintaining high computational efficiency. Emprical validation on a synthetic cosine dataset and real-world air quality sensor readings demonstrates performance competitive or superior to the state-of-the-art. By enhancing reliability beyond in-distribution scenarios, these results hold significant implications for the wider adoption of artificial intelligence in safety-critical applications.

泛化能力权重建模不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。