arXiv:2501.00999cs.CLcs.AI2025-01被引 9

用信息瓶颈理论揭示大模型如何压缩和提取信息

Exploring Information Processing in Large Language Models: Insights from Information Bottleneck Theory

  • 从信息瓶颈视角构建任务空间,分析模型压缩与选择信息的机制
  • 提出新方法使推理速度提升40%以上,且在多个数据集上表现更优
  • 适合研究大模型内部机理或优化推理效率的研究者

大语言模型在多种任务中表现出色,但其内部如何理解输入并做出有效预测仍不清晰。本文基于信息瓶颈理论,提出一种无需训练的任务空间构建策略,发现:大模型会将输入信息压缩至特定任务空间(如情感空间、主题空间)以促进任务理解,并在关键时刻从该空间中提取相关特征进行准确预测。基于此,提出两种新方法:基于信息压缩的上下文学习(IC-ICL),通过将检索示例压缩至任务空间提升推理效率;任务空间引导微调(TS-FT),采用空间引导损失鼓励模型学习更有效的压缩与选择机制。实验验证了任务空间构建的有效性。IC-ICL不仅提升性能,还使推理速度加快超过40%;TS-FT仅需微调调整即取得更优结果。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of tasks by understanding input information and predicting corresponding outputs. However, the internal mechanisms by which LLMs comprehend input and make effective predictions remain poorly understood. In this paper, we explore the working mechanism of LLMs in information processing from the perspective of Information Bottleneck Theory. We propose a non-training construction strategy to define a task space and identify the following key findings: (1) LLMs compress input information into specific task spaces (e.g., sentiment space, topic space) to facilitate task understanding; (2) they then extract and utilize relevant information from the task space at critical moments to generate accurate predictions. Based on these insights, we introduce two novel approaches: an Information Compression-based Context Learning (IC-ICL) and a Task-Space-guided Fine-Tuning (TS-FT). IC-ICL enhances reasoning performance and inference efficiency by compressing retrieved example information into the task space. TS-FT employs a space-guided loss to fine-tune LLMs, encouraging the learning of more effective compression and selection mechanisms. Experiments across multiple datasets validate the effectiveness of task space construction. Additionally, IC-ICL not only improves performance but also accelerates inference speed by over 40\%, while TS-FT achieves superior results with a minimal strategy adjustment.

大模型机理信息瓶颈推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。