通过模拟推理过程,解析Transformer在上下文线性回归中的测试时计算机制。
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- 用噪声注入和采样模拟语言模型解码过程
- 验证多种推理技术在真实模型中的有效性
- 为理解实际语言模型推理行为提供理论支持
在语言模型推理中增加测试时计算(如生成更多中间思考或采样多个候选答案)已被证明能显著提升性能。本文首次尝试弥合实践推理与理论Transformer分析之间的鸿沟,引入随机性和采样机制。研究聚焦于带有连续/二值系数的上下文线性回归任务,通过噪声注入和二值系数采样模拟语言模型的解码过程。基于该框架,对广泛采用的推理技术进行了详细分析。结合实证结果,本理论框架揭示了理解现实语言模型推理行为的新可能。
原文摘要 · Abstract (English)
Using more test-time computation during language model inference, such as generating more intermediate thoughts or sampling multiple candidate answers, has proven effective in significantly improving model performance. This paper takes an initial step toward bridging the gap between practical language model inference and theoretical transformer analysis by incorporating randomness and sampling. We focus on in-context linear regression with continuous/binary coefficients, where our framework simulates language model decoding through noise injection and binary coefficient sampling. Through this framework, we provide detailed analyses of widely adopted inference techniques. Supported by empirical results, our theoretical framework and analysis demonstrate the potential for offering new insights into understanding inference behaviors in real-world language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。