arXiv:2511.09381cs.CLcs.AI2025-11被引 4

对比大模型在生成与选择任务中的自修正能力差异。

Self-Correcting Large Language Models: Generation vs. Multiple Choice

  • 比较开放式生成与多选题两种任务的自修正机制
  • 生成任务更受益于灵活重解释,多选受限于选项边界
  • 为智能体系统设计提供任务结构适配依据

大型语言模型近期展现出通过迭代优化实现自我修正的能力,常被称为自一致性或自我反思。然而,这一自修正机制在开放式文本生成与从预定义选项中选择最优回答的任务中可能表现出显著不同的动态特征。本文系统比较了不同规模和架构的语言模型在多种自然语言理解与推理任务中,这两种范式下的性能趋势与纠错行为。实验结果揭示出截然不同的提升模式与失败形态:开放式生成通常得益于重新解读与组合式精炼带来的灵活性,而多选任务虽能利用更清晰的解空间边界,却可能受限于给出的选项范围。这一差异也反映了新兴智能体应用面临的双重挑战——既要能生成并优化开放式计划或解释,又需在受限动作空间中做出可靠离散决策。因此,本研究强调自修正机制的设计应考虑任务结构与输出空间之间的相互作用,对知识密集型推理与决策导向的应用具有重要意义。

原文摘要 · Abstract (English)

Large language models have recently demonstrated remarkable abilities to self-correct their responses through iterative refinement, often referred to as self-consistency or self-reflection. However, the dynamics of this self-correction mechanism may differ substantially depending on whether the model is tasked with open-ended text generation or with selecting the most appropriate response from multiple predefined options. In this paper, we conduct a systematic investigation of these two paradigms by comparing performance trends and error-correction behaviors across various natural language understanding and reasoning tasks, covering language models of different scales and families. Our experimental results reveal distinct patterns of improvement and failure modes: \textit{While open-ended generation often benefits from the flexibility of re-interpretation and compositional refinement, multiple-choice selection can leverage clearer solution boundaries but may be limited by the provided options}. This contrast also reflects the dual demands faced by emerging agentic LLM applications: effective agents must not only generate and refine open-ended plans or explanations, but also make reliable discrete choices when operating within constrained action spaces. Our findings, therefore, highlight that the design of self-correction mechanisms should take into account the interaction between task structure and output space, with implications for both knowledge-intensive reasoning and decision-oriented applications of LLMs.

自修正大模型推理任务设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。