arXiv:2510.22993cs.LGcs.CL2025-10

语言模型在上下文里组合基础技能完成复杂任务能力有限。

Can Language Models Compose Skills In-Context?

  • 用上下文示例引导模型组合基础技能完成复合任务
  • 简单示例反而降低性能,因模型难正确识别与组装技能
  • 对齐示例与步骤可显著提升表现,适合研究推理机制者

将基础技能从简单任务中组合以完成复合任务,是现代智能系统的关键。本文研究语言模型在上下文中的组合能力,即通过上下文示例展示基础技能来完成复合任务,这比训练中学习技能组合更具挑战性。我们在多个代表性开源语言模型上进行系统实验,采用语言和逻辑任务探测组合能力。结果表明,简单任务示例可能产生意外的负面影响,因模型普遍难以正确识别和组装技能,即使使用思维链示例亦然。理论分析进一步显示,示例与组合步骤的对齐至关重要。据此提出一种探测任务方法,其性能提升为该洞察提供了积极支持。

原文摘要 · Abstract (English)

Composing basic skills from simple tasks to accomplish composite tasks is crucial for modern intelligent systems. We investigate the in-context composition ability of language models to perform composite tasks that combine basic skills demonstrated in in-context examples. This is more challenging than the standard setting, where skills and their composition can be learned in training. We conduct systematic experiments on various representative open-source language models, utilizing linguistic and logical tasks designed to probe composition abilities. The results reveal that simple task examples can have a surprising negative impact on the performance, because the models generally struggle to recognize and assemble the skills correctly, even with Chain-of-Thought examples. Theoretical analysis further shows that it is crucial to align examples with the corresponding steps in the composition. This inspires a method for the probing tasks, whose improved performance provides positive support for our insights.

语言模型技能组合推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。