arXiv:2604.22770cs.CYcs.AI2026-04

用多智能体辩论评估对话能力,实现精准个性化语言学习进度

Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning

论文配图:Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning
图 1 · 摘自论文原文
  • 多智能体分角色独立评分后辩论,达成共识得分
  • 90.91%推荐接受度,70%掌握度才可升级
  • 适合需要真实对话提升的学习者

现有数字语言学习多依赖碎片化测验,仅检验记忆而非实际对话能力。当进度由测验成绩决定时,学习者可能在互动中仍存在语法与词汇使用缺陷。我们提出 Learning in Blocks 框架,基于 CEFR 标准评估对话表现。该框架采用异构多智能体辩论(HeteroMAD)分两阶段:评分阶段中,分工角色的智能体分别评估语法、词汇与交互沟通,并通过辩论解决分歧;判断者综合得出一致评分;推荐阶段则定位需强化的语法与词汇点。进度要求达到 70% 掌握度,且通过间隔复习巩固薄弱项以防止技能退化。我们在 ESL 专家标注的 CEFR A2 对话数据集上测试四类方法,HeteroMAD 与人工评分一致性达 0.23 的偏差,推荐可接受度为 90.91%。为期八周、覆盖 180 名 A2 学习者的实验表明,结合标准评分、个性化推荐、间隔复习与掌握度推进,效果优于仅提供反馈。

原文摘要 · Abstract (English)

Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. When progression is driven by quiz performance, learners can advance despite persistent gaps in using grammar and vocabulary during interaction. Recent work on LLM-based judging suggests a path toward scoring open-ended conversations, but using interaction evidence to drive progression and review requires scoring protocols that are reliable and validated. We introduce Learning in Blocks, a framework that grounds progression in demonstrated conversational competence evaluated using CEFR-aligned rubrics. The framework employs heterogeneous multi-agent debate (HeteroMAD) in two stages: a scoring stage where role-specialized agents independently evaluate Grammar, Vocabulary, and Interactive Communication, engage in debate to address conflicting judgments, and a judge synthesizes consensus scores; and a recommendation stage that identifies specific grammar skills and vocabulary topics for targeted review. Progression requires demonstrating 70% mastery, and spaced review targets identified weaknesses to counter skill decay. We benchmark four scoring and recommendation methods on CEFR A2 conversations annotated by ESL experts. HeteroMAD achieves a superior score agreement with a 0.23 degree of variation and recommendation acceptability of 90.91%. An 8-week study with 180 CEFR A2 learners demonstrates that combining rubric-aligned scoring and recommendation with spaced review and mastery-based progression produces better learning outcomes than feedback alone.

语言学习多智能体自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。