arXiv:2502.11258cs.CL2025-02被引 2

用信息论优化大模型微调,提升分类性能。

Leveraging Conditional Mutual Information to Improve Large Language Model Fine-Tuning For Classification

  • 通过最小化条件互信息提升大模型独立性能。
  • 在6/8 GLUE任务上优于BERT,知识蒸馏效果更优。
  • 适合关注模型优化与信息论融合的研究者。

尽管大语言模型近年来展现出卓越能力,但信息论在大模型发展中的潜力仍待挖掘。本文将条件互信息(CMI)引入大模型分类任务的微调中,探索其两种应用:一是通过最小化CMI提升模型独立性能,二是通过最大化CMI增强知识蒸馏(KD),训练出更强的学生模型。为在大模型微调中应用CMI,我们对近期提出的约束条件互信息的深度学习框架进行适配与修改,该框架最初用于图像分类。实验表明,在6/8个GLUE分类任务上,最小化CMI的微调策略优于BERT;在知识蒸馏过程中,最大化CMI使学生模型在6/8个任务上超越DistilBERT。结果证明,CMI可有效优化独立大模型与学生模型,展现其作为大模型微调稳健框架的潜力。本研究弥合了信息论与大模型发展的鸿沟,为构建高性能语言模型提供新思路。

原文摘要 · Abstract (English)

Although large language models (LLMs) have demonstrated remarkable capabilities in recent years, the potential of information theory (IT) to enhance LLM development remains underexplored. This paper introduces the information theoretic principle of Conditional Mutual Information (CMI) to LLM fine-tuning for classification tasks, exploring its promise in two main ways: minimizing CMI to improve a model's standalone performance and maximizing CMI to enhance knowledge distillation (KD) for more capable student models. To apply CMI in LLM fine-tuning, we adapt the recently proposed CMI-constrained deep learning framework, which was initially developed for image classification, with some modification. By minimizing CMI during LLM fine-tuning, we achieve superior performance gains on 6 of 8 GLUE classification tasks compared to BERT. Additionally, maximizing CMI during the KD process results in significant performance improvements in 6 of 8 GLUE classification tasks compared to DistilBERT. These findings demonstrate CMI's adaptability for optimizing both standalone LLMs and student models, showcasing its potential as a robust framework for advancing LLM fine-tuning. Our work bridges the gap between information theory and LLM development, offering new insights for building high-performing language models.

信息论微调优化知识蒸馏大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。