arXiv:2506.04690cs.LGcs.AI2025-06被引 1

通过输入分布投影提升模型泛化能力

Towards Better Generalization via Distributional Input Projection Network

  • 在每层将输入映射为可学习的分布,使损失曲面更平滑
  • 实验证明在视觉、语言模型上均提升测试性能与抗干扰能力
  • 方法通用易用,适合主流深度学习模型

随着过参数化模型日益普及,仅靠训练损失已难以反映泛化性能。尽管平滑性已被证明有助于提升泛化,但直接在神经网络中强制平滑仍具挑战。为此,我们提出分布式输入投影网络(DIPNet),在每一层将输入投影到可学习的分布空间。该分布表示使输入相关的损失曲面更加平滑,从而促进更好泛化。理论分析表明,DIPNet降低了局部平滑度量和网络的Lipschitz常数。实验验证了其在多种架构和任务中的有效性,包括视觉变换器(ViTs)、大语言模型(LLMs)、ResNet和MLPs。该方法在标准设置、对抗攻击、分布外输入及推理基准上均显著提升测试性能。结果表明,输入投影策略可无缝集成至现有模型,为现代深度学习提供一种通用且高效的泛化增强方案。

原文摘要 · Abstract (English)

As overparameterized models become increasingly prevalent, training loss alone offers limited insight into generalization performance. While smoothness has been linked to improved generalization across various settings, directly enforcing smoothness in neural networks remains challenging. To address this, we introduce Distributional Input Projection Networks (DIPNet), a novel framework that projects inputs into learnable distributions at each layer. This distributional representation induces a smoother loss landscape with respect to the input, promoting better generalization. We provide theoretical analysis showing that DIPNet reduces both local smoothness measures and the Lipschitz constant of the network, contributing to improved generalization performance. Empirically, we validate DIPNet across a wide range of architectures and tasks, including Vision Transformers (ViTs), Large Language Models (LLMs), ResNet and MLPs. Our method consistently enhances test performance under standard settings, adversarial attacks, out-of-distribution inputs, and reasoning benchmarks. We demonstrate that the proposed input projection strategy can be seamlessly integrated into existing models, providing a general and effective approach for boosting generalization performance in modern deep learning.

泛化能力输入投影平滑性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。