arXiv:2503.11336cs.CL2025-03被引 3

用规则引导反馈,让大模型更准确地遵守指令并主动查漏补缺。

Rule-Guided Feedback: Enhancing Reasoning by Enforcing Rule Adherence in Large Language Models

  • 通过教师模型严格检查学生输出是否符合任务规则,提供改进建议而非直接答案。
  • 在多个任务上显著提升大模型表现,如国际象棋残局和数学推理。
  • 适合需要高可靠性和规则约束的场景,如自动评测与智能辅导系统。

本文提出规则引导反馈(Rule-Guided Feedback, RGF)框架,通过结构化规则遵循和主动信息获取来提升大语言模型(LLM)性能。RGF采用师生范式,由教师模型严格评估学生输出是否符合任务特定规则,在发现偏差时提供建设性指导而非直接答案。这种迭代反馈循环兼具双重作用:确保解题过程符合预设约束,并激励模型主动寻求缺失信息以解决不确定性。我们在多样任务上评估RGF,包括一招将死谜题、十四行诗创作、企鹅桌位分类、GSM8k和StrategyQA。结果表明,结构化反馈机制可显著提升大模型在多领域中的表现。

原文摘要 · Abstract (English)

In this paper, we introduce Rule-Guided Feedback (RGF), a framework designed to enhance Large Language Model (LLM) performance through structured rule adherence and strategic information seeking. RGF implements a teacher-student paradigm where rule-following is forced through established guidelines. Our framework employs a Teacher model that rigorously evaluates each student output against task-specific rules, providing constructive guidance rather than direct answers when detecting deviations. This iterative feedback loop serves two crucial purposes: maintaining solutions within defined constraints and encouraging proactive information seeking to resolve uncertainties. We evaluate RGF on diverse tasks including Checkmate-in-One puzzles, Sonnet Writing, Penguins-In-a-Table classification, GSM8k, and StrategyQA. Our findings suggest that structured feedback mechanisms can significantly enhance LLMs' performance across various domains.

规则遵循反馈机制大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。