arXiv:2603.27076cs.AI2026-03被引 1

验证反馈在逻辑证明辅导中可能适得其反,需根据初始反馈质量动态调整。

When Verification Hurts: Asymmetric Effects of Multi-Agent Feedback in Logic Proof Tutoring

  • 基于知识图谱构建516个证明状态的细粒度标注数据集
  • 验证器在错误反馈时提升性能,但在可靠反馈时反而降低4-6个百分点
  • 所有模型在复杂度4-5以上均难以突破,存在共性瓶颈

大语言模型在自动化辅导中的可靠性在符号推理领域仍不明确。本文研究命题逻辑证明的步骤级反馈,要求与学习者当前证明状态精确对齐。我们构建了一个基于知识图谱的基准数据集,包含516个独特证明状态,附有步骤级标注和难度指标。不同于以往依赖模型自评或二元正确性的评估方式,本框架可针对已验证解题路径进行细粒度反馈质量分析。评估三种角色专用流程:仅部分解题路径可见的Tutor、全推导路径可见的Teacher,以及验证Tutor反馈的Judge。结果揭示显著不对称性:当上游反馈错误率较高(<70%准确率)时,验证能提升效果;但当反馈已较可靠(>85%)时,验证会导致4-6个百分点的性能下降,源于过度指定。关键发现是存在共享复杂度上限——任何模型或流程在复杂度4-5以上的证明状态上均无法稳定成功。这挑战了‘添加验证器或更丰富上下文总能改进辅导’的假设,呼吁采用适应性、难度感知的架构,按估计复杂度和上游可靠性动态路由问题。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for automated tutoring, but their reliability in structured symbolic domains remains unclear. We study step-level feedback for propositional logic proofs, which require precise symbolic reasoning aligned with a learner's current proof state. We introduce a knowledge-graph-grounded benchmark of 516 unique proof states with step-level annotations and difficulty metrics. Unlike prior tutoring evaluations that rely on model self-assessment or binary correctness, our framework enables fine-grained analysis of feedback quality against verified solution paths. We evaluate three role-specialized pipelines with varying solution access: Tutor (partial solution access), Teacher (full derivation access), and Judge (verification of Tutor feedback). Our results reveal a striking asymmetry: verification improves outcomes when upstream feedback is error-prone (<70% accuracy), but degrades performance by 4-6 percentage points through over-specification when feedback is already reliable (>85%). Critically, we identify a shared complexity ceiling; no model or pipeline reliably succeeds on proof states exceeding complexity 4-5. These findings challenge the assumption that adding verifiers or richer context universally improves tutoring, motivating adaptive, difficulty-aware architectures that route problems by estimated complexity and upstream reliability.

逻辑推理辅导系统LLM评估多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。