arXiv:2503.00902cs.CL2025-03NAACL被引 9

动态指令让大模型自我反思更准更快,数学常识题表现提升10.1%。

Instruct-of-Reflection: Enhancing Large Language Models Iterative Reflection Capabilities via Dynamic-Meta Instruction

  • 用动态元指令引导反思,避免固定套路导致的重复和偏差。
  • 在数学与常识推理任务上,平均性能比基线高10.1%。
  • 适合需要持续优化输出质量的对话、写作等场景。

大语言模型的自反思机制受到广泛关注。现有方法依赖模型内部自我修正或外部反馈来迭代改进回答,但近期研究质疑纯内生自纠错可能反而降低性能。我们实证发现,当前静态反思方法易引发冗余、漂移与固执问题。为此提出Instruct-of-Reflection(IoRT),一种新颖通用的反思框架,通过动态元指令增强模型的迭代反思能力。具体地,由元思维与自一致性分类器驱动的导师生成包括刷新、停止、选择在内的多种指令,指导下一轮反思。实验表明,IoRT在数学与常识推理任务上相较成熟基线平均提升10.1%,验证了其有效性与适用性。

原文摘要 · Abstract (English)

Self-reflection for Large Language Models (LLMs) has gained significant attention. Existing approaches involve models iterating and improving their previous responses based on LLMs' internal reflection ability or external feedback. However, recent research has raised doubts about whether intrinsic self-correction without external feedback may even degrade performance. Based on our empirical evidence, we find that current static reflection methods may lead to redundant, drift, and stubborn issues. To mitigate this, we introduce Instruct-of-Reflection (IoRT), a novel and general reflection framework that leverages dynamic-meta instruction to enhance the iterative reflection capability of LLMs. Specifically, we propose the instructor driven by the meta-thoughts and self-consistency classifier, generates various instructions, including refresh, stop, and select, to guide the next reflection iteration. Our experiments demonstrate that IoRT achieves an average improvement of 10.1% over established baselines in mathematical and commonsense reasoning tasks, highlighting its efficacy and applicability.

大模型自我反思动态指令推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。