arXiv:2507.12004cs.CL2025-07

通过分析模型表征,提升语言模型的数据与参数效率。

Improving Data and Parameter Efficiency of Neural Language Models Using Representation Analysis

  • 基于表征平滑性设计正则化策略,稳定训练过程。
  • 结合主动学习与参数高效微调,减少标注需求和计算资源。
  • 适合低资源场景下追求高效训练的研究者与开发者。

本论文针对神经语言模型在数据与参数效率方面的挑战,聚焦表征分析与新型优化技术。第一部分研究语言模型内部表征的特性与动态,提出基于表征平滑性的创新方法,利用雅可比与海森矩阵设计正则化策略,增强模型对输入扰动的鲁棒性。第二部分结合表征平滑性洞察,将主动学习与参数高效微调融合,提出平滑性指导的早停机制,无需标注验证集;并设计新组合策略,显著降低标注成本与计算开销。第三部分探索以上下文学习为弱监督机制,有效利用未标注数据,在低资源与动态数据环境下提升模型泛化能力。大量实验证明,该系列方法在性能、稳定性与效率上均显著优于传统方法。

原文摘要 · Abstract (English)

This thesis addresses challenges related to data and parameter efficiency in neural language models, with a focus on representation analysis and the introduction of new optimization techniques. The first part examines the properties and dynamics of language representations within neural models, emphasizing their significance in enhancing robustness and generalization. It proposes innovative approaches based on representation smoothness, including regularization strategies that utilize Jacobian and Hessian matrices to stabilize training and mitigate sensitivity to input perturbations. The second part focuses on methods to significantly enhance data and parameter efficiency by integrating active learning strategies with parameter-efficient fine-tuning, guided by insights from representation smoothness analysis. It presents smoothness-informed early-stopping techniques designed to eliminate the need for labeled validation sets and proposes innovative combinations of active learning and parameter-efficient fine-tuning to reduce labeling efforts and computational resources. Extensive experimental evaluations across various NLP tasks demonstrate that these combined approaches substantially outperform traditional methods in terms of performance, stability, and efficiency. The third part explores weak supervision techniques enhanced by in-context learning to effectively utilize unlabeled data, further reducing dependence on extensive labeling. It shows that using in-context learning as a mechanism for weak supervision enables models to better generalize from limited labeled data by leveraging unlabeled examples more effectively during training. Comprehensive empirical evaluations confirm significant gains in model accuracy, adaptability, and robustness, especially in low-resource settings and dynamic data environments.

表征分析参数效率主动学习弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。