arXiv:2605.13511cs.CLcs.AI2026-05中稿 · ICML被引 3

让大模型通过示范学习推理,提升表现。

Many-Shot CoT-ICL: Making In-Context Learning Truly Learn

论文配图:Many-Shot CoT-ICL: Making In-Context Learning Truly Learn
图 1 · 摘自论文原文
  • 将示范按认知难度排序,构建渐进式学习路径。
  • 在数学任务上,性能提升最高达5.42个百分点。
  • 适合研究大模型推理能力与提示工程的学者。

尽管多示范上下文学习(many-shot ICL)表现出色,但以往研究主要集中在非推理任务。本文首次系统分析了多示范思维链上下文学习(CoT-ICL)在推理任务中的缩放行为。通过对非推理与推理任务、以及非推理和推理导向的大语言模型进行对比,我们发现多示范CoT-ICL具有若干独特性质。我们将其视为测试时的上下文学习而非简单的模式匹配,并提出两条原则:(i) 示范应易于目标模型理解;(ii) 示范应按支持平滑概念演进的方式排序。基于此,我们提出曲线示范选择(CDS)方法,仅通过简单排序,在64个示范的数学任务上实现最高5.42个百分点的性能提升。整体而言,本工作将长上下文窗口从检索缓冲区重构为结构化的测试时学习课程。

原文摘要 · Abstract (English)

While many-shot ICL achieves remarkable performance, prior studies of its scaling behavior have mainly focused on non-reasoning tasks. In this work, we study many-shot ICL on reasoning tasks, with a particular focus on many-shot chain-of-thought in-context learning (CoT-ICL). Analyzing across non-reasoning and reasoning tasks and across non-reasoning and reasoning-oriented LLMs, we identify several distinctive properties of many-shot CoT-ICL. We further interpret these findings by viewing many-shot CoT-ICL as in-context test-time learning rather than scaled pattern matching, and suggest two principles: (i) demonstrations should be easy for the target model to understand, and (ii) they should be ordered to support a smooth conceptual progression. Guided by the principle, we propose Curvilinear Demonstration Selection (CDS), a simple ordering method that yields up to a 5.42 percentage-point gain on a math task with 64 demonstrations. Overall, our results reframe the long context window from a retrieval buffer into a structured curriculum for in-context test-time learning.

上下文学习思维链大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。