用信息瓶颈理论解释双向模型为何更优
FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models
- 提出新方法FlowNIB动态估算训练中互信息
- 发现双向模型保留更多有效信息且维度更高
- 适合研究模型表征与信息压缩的学者
双向语言模型在自然语言理解任务中表现优于单向模型,但其理论原因尚不明确。本文基于信息瓶颈(IB)原理,研究这一差异。提出一种可动态、可扩展估计训练中互信息的新方法FlowNIB,克服了传统IB方法计算不可行和固定权衡率的缺陷。理论上证明,双向模型保留的互信息更多,有效维度更高;并建立广义表征复杂度测量框架,证明在温和条件下双向表示严格更丰富。通过在多个模型和任务上的大量实验验证,揭示了信息在训练过程中的编码与压缩机制。本工作为双向架构的有效性提供了理论依据,并提供了一个分析深度语言模型信息流的实用工具。
原文摘要 · Abstract (English)
Bidirectional language models have better context understanding and perform better than unidirectional models on natural language understanding tasks, yet the theoretical reasons behind this advantage remain unclear. In this work, we investigate this disparity through the lens of the Information Bottleneck (IB) principle, which formalizes a trade-off between compressing input information and preserving task-relevant content. We propose FlowNIB, a dynamic and scalable method for estimating mutual information during training that addresses key limitations of classical IB approaches, including computational intractability and fixed trade-off schedules. Theoretically, we show that bidirectional models retain more mutual information and exhibit higher effective dimensionality than unidirectional models. To support this, we present a generalized framework for measuring representational complexity and prove that bidirectional representations are strictly more informative under mild conditions. We further validate our findings through extensive experiments across multiple models and tasks using FlowNIB, revealing how information is encoded and compressed throughout training. Together, our work provides a principled explanation for the effectiveness of bidirectional architectures and introduces a practical tool for analyzing information flow in deep language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。