大模型能在一次推理中同时执行多个任务,展现惊人并行能力。
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
- 模型在单次推理中可并行处理多个不同任务,无需切换。
- 模型规模越大,并行处理任务越多且输出更精准。
- 揭示了大模型作为多重模拟器的潜在机制,适合研究者深挖。
大型语言模型(LLMs)展现出强大的上下文学习(ICL)能力。本研究探索了一个令人惊讶的现象:LLMs可在单次推理中同时执行多个计算上不同的ICL任务,我们称之为“任务超叠加”。我们在多种模型家族和规模下提供了实证证据,并发现即使模型仅训练为一次学习一个任务,该现象仍会浮现。我们从理论上解释,这一能力完全在Transformer的表达能力范围内。此外,我们分析了模型在超叠加过程中如何内部组合任务向量。还发现更大的模型能并行解决更多任务,且更优地校准输出分布。这些发现揭示了大模型的潜在能力,进一步支持了“大模型即多重模拟器叠加”的观点,并引发对并发任务执行机制的思考。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable in-context learning (ICL) capabilities. In this study, we explore a surprising phenomenon related to ICL: LLMs can perform multiple, computationally distinct ICL tasks simultaneously, during a single inference call, a capability we term "task superposition". We provide empirical evidence of this phenomenon across various LLM families and scales and show that this phenomenon emerges even if we train the model to in-context learn one task at a time. We offer theoretical explanations that this capability is well within the expressive power of transformers. We also explore how LLMs internally compose task vectors during superposition. Furthermore, we show that larger models can solve more ICL tasks in parallel, and better calibrate their output distribution. Our findings offer insights into the latent capabilities of LLMs, further substantiate the perspective of "LLMs as superposition of simulators", and raise questions about the mechanisms enabling simultaneous task execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。