arXiv:2601.11979cs.AIcs.LG2026-01被引 1

动态插入示范案例,提升大模型数学推理准确率

Process In-Context Learning: Enhancing Mathematical Reasoning via Dynamic Demonstration Insertion

  • 根据推理过程实时识别困惑点并提取特征
  • 在推理中动态插入匹配示范,减少逻辑错误
  • 适合需要逐步推导的数学题与复杂推理任务

上下文学习(ICL)在多种大语言模型任务中表现优异,但在需要逐步逻辑推导的任务(如数学推理)中潜力尚未充分挖掘。现有方法使用固定示范,无法随推理过程动态调整,导致计算歧义或逻辑断层等困惑点难以解决,进而引发连锁错误。为此,我们提出过程式上下文学习(PICL),一种动态示范集成框架。PICL分两阶段运行:1)通过分析推理过程中的语义与熵值识别潜在困惑点,并提取其核心特征;2)当遇到这些困惑点时,从示范池中检索匹配上下文的示范并直接插入当前推理流程,引导后续步骤。实验表明,PICL通过缓解推理中段的困惑问题,显著优于基线方法,验证了动态示范插入在复杂数学推理中的有效性。

原文摘要 · Abstract (English)

In-context learning (ICL) has proven highly effective across diverse large language model (LLM) tasks. However, its potential for enhancing tasks that demand step-by-step logical deduction, such as mathematical reasoning, remains underexplored. A core limitation of existing ICL approaches is their static use of demonstrations: examples are pre-selected before inference and remain fixed, failing to adapt to the dynamic confusion points that often arise during multi-step reasoning such as ambiguous calculations or logical gaps. These unresolved confusion points can lead to cascading errors that degrade final accuracy. To tackle this issue, we propose Process In-Context Learning (PICL), a dynamic demonstration integration framework designed to boost mathematical reasoning by responding to real-time inference needs. PICL operates in two stages: 1)~it identifies potential confusion points by analyzing semantics and entropy in the reasoning process and summarizes their core characteristics; 2)~upon encountering these points, it retrieves relevant demonstrations from the demonstration pool that match the confusion context and inserts them directly into the ongoing reasoning process to guide subsequent steps. Experiments show that PICL outperforms baseline methods by mitigating mid-inference confusion, highlighting the value of adaptive demonstration insertion in complex mathematical reasoning.

数学推理动态示范大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。