arXiv:2606.23668cs.LG2026-06

语言模型提示学习有本质局限,无法解决所有任务。

On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners

  • 将用户-系统交互建模为廉价对话博弈,分析提示中的任务编码与重构机制。
  • 证明存在不可消除的错误下限,即使无限数据也无法突破。
  • 适合关注大模型能力边界、提示工程局限的研究者阅读。

大型语言模型(LLMs)常被视为通用求解器,能够应对任意任务。本文指出,这种观点忽视了一个根本限制:语言是传递任务信息的压缩且容量受限的接口。我们将用户-系统交互建模为双层‘廉价对话’博弈,分析潜藏任务如何被编码进提示,并在对齐与安全约束下被重新解释。我们提出概念分解,区分任务推断与执行,并推导出PAC-Bayes界,以区分有限样本估计误差与不可消除的结构性限制。第一个主要结果确立了‘表达力下限’:语言作为容量受限的通信信道,当任务族的信息复杂度超过信道容量时,不同任务不可避免地变得不可区分,导致严格正的误差下限,无法仅通过更多数据、优化或模型扩展消除。第二个结果建立‘目标不匹配下限’:当对齐约束限制可接受输出集时,用户理想分布可能不在可行类中,引发不可消除的失真。二者共同得出一个形式化否定结论:仅靠提示,语言模型无法成为通用问题求解器,因为存在某些任务族,其正确行为在无限数据条件下仍不可实现。更广泛地,我们的分析表明,基于提示的泛化极限源于信息受限的通信和对齐约束的目标。这暗示,超越自然语言的接口,如多模态观察和外部记忆,可能通过增加系统可用的任务相关信息,缓解大模型的固有局限。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are frequently portrayed as general-purpose solvers capable of solving arbitrary tasks. We argue that this view overlooks a fundamental constraint: language is a compressed and capacity-limited interface for conveying task information. Modelling User--System interaction as a bilevel \emph{cheap-talk} game, we analyse how latent tasks are encoded into prompts and reinterpreted under alignment and safety constraints. We introduce a conceptual decomposition separating task inference from execution and derive PAC-Bayes bounds that distinguish finite-sample estimation error from irreducible structural limitations. Our first main result establishes an \emph{expressivity floor}: language acts as a capacity-limited communication channel, and whenever the informational complexity of a task family exceeds the capacity of that channel, distinct tasks become unavoidably indistinguishable to the Solver, inducing a strictly positive error floor that cannot be eliminated by additional data, optimisation, or model scaling alone. We then establish an \emph{objective-misalignment floor}: when alignment constraints restrict the admissible output set, the User-ideal distribution may lie outside the feasible class, inducing an irreducible distortion. Together, these results yield a formal negative conclusion: prompt-conditioned LLMs are not universal problem solvers through prompting alone, as there exist task families for which correct behaviour is provably unattainable even in the infinite-data regime. More broadly, our analysis shows the limits of prompt-based generalisation arise from information-constrained communication and alignment-constrained objectives. This suggests that interfaces beyond natural language, including multimodal observations and, external memory, may reduce the inherent LLM limitations by increasing the task-relevant information available to the System.

大模型局限提示工程信息瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。