arXiv:2502.13311cs.CLcs.AI2025-02ACL被引 22

用逐轮验证机制提升大模型编程导师的引导能力

Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors

  • 设计追踪-验证流程,实时评估学生知识状态并逐步检查代码
  • 在模拟学生测试中,新方法成功率显著高于传统方式
  • 适合研究智能辅导系统或大模型应用的开发者参考

基于大语言模型(LLMs)的智能辅导代理在语言学习和科学教育中已得到广泛应用,但在指导用户完成复杂现实任务方面仍显不足。为此,本文聚焦编程辅导这一高难度任务,提出一种新型代理工作流——追踪-验证(TRAVER),结合知识追踪以估计学生知识状态,并通过逐轮验证确保任务导向的有效引导。我们引入DICT自动评估协议,利用受控学生模拟与代码生成测试来评估导师代理性能。大量实验揭示了编程辅导中的关键挑战,并证明TRAVER在任务完成率上显著更优。尽管本文以编程辅导为例,但该方法可拓展至其他领域,为人类任务学习中的辅导代理发展提供重要启示。

原文摘要 · Abstract (English)

Intelligent tutoring agents powered by large language models (LLMs) have been increasingly explored to deliver personalized knowledge in areas such as language learning and science education. However, their capabilities in guiding users to solve complex real-world tasks remain underexplored. To address this limitation, in this work, we focus on coding tutoring, a challenging problem that requires tutors to proactively guide students towards completing predefined coding tasks. We propose a novel agent workflow, Trace-and-Verify (TRAVER), which combines knowledge tracing to estimate a student's knowledge state and turn-by-turn verification to ensure effective guidance toward task completion. We introduce DICT, an automatic evaluation protocol that assesses tutor agents using controlled student simulation and code generation tests. Extensive experiments reveal the challenges of coding tutoring and demonstrate that TRAVER achieves a significantly higher success rate. Although we use code tutoring as an example in this paper, our approach can be extended beyond coding, providing valuable insights into advancing tutoring agents for human task learning.

智能辅导大模型应用编程教育对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。