arXiv:2506.01034cs.CLcs.AI2025-06中稿 · NeurIPS

通过分析嵌入空间局部维度,揭示微调对大模型的影响规律。

Less is More: Local Intrinsic Dimensions of Contextual Language Models

  • 用嵌入空间局部维度衡量模型训练状态
  • 维度下降预示性能提升,反映过拟合或模型觉醒
  • 为微调策略提供可量化的几何洞察

理解大语言模型(LLMs)的内部机制仍具挑战性,许多基础问题如微调如何影响模型行为,往往需大量实证评估。本文基于上下文嵌入的几何特性,提出新视角:测量上下文语言模型隐空间的局部维度,并分析其在训练和微调过程中的变化。结果表明,局部维度均值能预测模型训练能力耗尽、过拟合及模型觉醒(grokking)现象。例如,在对话状态追踪任务中体现训练极限,在情感识别任务中显示过拟合,在算术任务中呈现模型觉醒。实验还发现,局部维度均值的下降常伴随且预示后续性能提升。该研究为从业者提供了微调对嵌入空间影响的深层理解,助力针对特定应用配置模型。成果推动了对大模型可解释性、适应性与泛化能力的讨论,弥合了内在机制与嵌入几何属性之间的鸿沟。

原文摘要 · Abstract (English)

Understanding the internal mechanisms of large language models (LLMs) remains a challenging and complex endeavor. Even fundamental questions, such as how fine-tuning affects model behavior, often require extensive empirical evaluation. In this paper, we introduce a novel perspective based on the geometric properties of contextual latent embeddings to study the effects of training and fine-tuning. To that end, we measure the local dimensions of a contextual language model's latent space and analyze their shifts during training and fine-tuning. We show that the local dimensions provide insights into the model's training dynamics and generalization ability. Specifically, the mean of the local dimensions predicts when the model's training capabilities are exhausted, as exemplified in a dialogue state tracking task, overfitting, as demonstrated in an emotion recognition task, and grokking, as illustrated with an arithmetic task. Furthermore, our experiments suggest a practical heuristic: reductions in the mean local dimension tend to accompany and predict subsequent performance gains. Through this exploration, we aim to provide practitioners with a deeper understanding of the implications of fine-tuning on embedding spaces, facilitating informed decisions when configuring models for specific applications. The results of this work contribute to the ongoing discourse on the interpretability, adaptability, and generalizability of LLMs by bridging the gap between intrinsic model mechanisms and geometric properties in the respective embeddings.

大模型机制嵌入空间微调分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。