arXiv:2601.06788cs.LGcs.AI2026-01被引 1

用量子纠缠视角揭示大模型微调中参数结构的内在规律

Artificial Entanglement in the Fine-Tuning of Large Language Models

  • 将低秩微调参数视为矩阵乘积态,量化其人工纠缠度
  • 发现LoRA更新存在中心抑制的体积律纠缠,而注意力层遵循面积律
  • 提出'无毛定理'类比:内部差异不影响输出效果,解释低秩方法高效原因

大语言模型可通过参数高效微调(PEFT)方法仅更新少量参数实现任务适配,通常采用低秩更新。本文从量子信息视角理解其有效性:低秩参数化自然对应于低维矩阵乘积态(MPS)表示,可对参数结构进行纠缠表征。我们定义并测量神经网络参数的人工纠缠熵。基于LLaMA模型在1B和8B规模上于Tulu3与OpenThoughts3数据集上的实验,发现:(i) LoRA中查询与值投影矩阵更新的内部人工纠缠遵循体积律且具中心抑制(称作'纠缠谷'),受超参数敏感,与全量微调(FFT)显著不同;(ii) 注意力矩阵的外部人工纠缠(对应表示空间中的词元-词元相关性)遵循面积律并带对数修正,对LoRA超参数和训练步数具有鲁棒性。类比黑洞物理的无毛定理,我们提出尽管LoRA与FFT在内部纠缠特征上存在差异,但其注意力输出并无明显区别,表明低秩更新具有'无毛'特性,解释了其有效性。我们进一步基于随机矩阵理论提供理论支持,并将分析扩展至MPS适配型PEFT方法,其行为定性相似。

原文摘要 · Abstract (English)

Large language models (LLMs) can be adapted to new tasks using parameter-efficient fine-tuning (PEFT) methods that modify only a small number of trainable parameters, often through low-rank updates. In this work, we adopt a quantum-information-inspired perspective to understand their effectiveness. From this perspective, low-rank parameterizations naturally correspond to low-dimensional Matrix Product States (MPS) representations, which enable entanglement-based characterizations of parameter structure. Thereby, we term and measure "Artificial Entanglement", defined as the entanglement entropy of the parameters in artificial neural networks (in particular the LLMs). We first study the representative low-rank adaptation (LoRA) PEFT method, alongside full fine-tuning (FFT), using LLaMA models at the 1B and 8B scales trained on the Tulu3 and OpenThoughts3 datasets, and uncover: (i) Internal artificial entanglement in the updates of query and value projection matrices in LoRA follows a volume law with a central suppression (termed as the "Entanglement Valley"), which is sensitive to hyper-parameters and is distinct from that in FFT; (ii) External artificial entanglement in attention matrices, corresponding to token-token correlations in representation space, follows an area law with logarithmic corrections and remains robust to LoRA hyper-parameters and training steps. Drawing a parallel to the No-Hair Theorem in black hole physics, we propose that although LoRA and FFT induce distinct internal entanglement signatures, such differences do not manifest in the attention outputs, suggesting a "no-hair" property that results in the effectiveness of low rank updates. We further provide theoretical support based on random matrix theory, and extend our analysis to an MPS Adaptation PEFT method, which exhibits qualitatively similar behaviors.

大模型微调量子信息低秩更新纠缠熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。