arXiv:2603.13353cs.AI2026-03中稿 · presentation at th…被引 1

用多智能体协作提升大模型对课堂对话标注的准确性与可靠性。

Optimizing LLM Annotation of Classroom Discourse through Multi-Agent Orchestration

  • 分三阶段:单次标注→自我验证→分歧仲裁,模拟人工标注流程。
  • 相比单次标注,关键教育概念标注准确率显著提升。
  • 适合需要高可靠标注的教育数据研究者使用。

大型语言模型(LLMs)正被广泛用于教育数据标注,如课堂对话、互动日志和学习成果。其快速生成教学互动摘要并匹配评分标准的能力,有望大幅降低专家人工标注的成本与时间。然而,现有证据表明,单一输出的LLM在需上下文、教学法或规范判断的高风险教育概念(如教学意图或对话行为)上仍不可靠。本文提出并实证评估了一种分层、成本感知的多智能体协同标注框架,通过建模计算权衡提升标注可靠性。该框架将标注视为多阶段认知过程:(1) 初步独立标注阶段,模型基于评分标准分配标签;(2) 自我验证阶段,模型依据评分定义审查自身输出并修正不一致之处;(3) 分歧仲裁阶段,独立仲裁模型审查已验证的标签与理由,最终确定符合评分标准的标签。该结构模仿了教育研究中常见的初始编码、自查与专家调解流程。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly positioned as scalable tools for annotating educational data, including classroom discourse, interaction logs, and qualitative learning artifacts. Their ability to rapidly summarize instructional interactions and assign rubric-aligned labels has fueled optimism about reducing the cost and time associated with expert human annotation. However, growing evidence suggests that single-pass LLM outputs remain unreliable for high-stakes educational constructs that require contextual, pedagogical, or normative judgment, such as instructional intent or discourse moves. This tension between scale and validity sits at the core of contemporary education data science. In this work, we present and empirically evaluate a hierarchical, cost-aware orchestration framework for LLM-based annotation that improves reliability while explicitly modeling computational tradeoffs. Rather than treating annotation as a one-shot prediction problem, we conceptualize it as a multi-stage epistemic process comprising (1) an unverified single-pass annotation stage, in which models independently assign labels based on the rubric; (2) a self-verification stage, in which each model audits its own output against rubric definitions and revises its label if inconsistencies are detected; and (3) a disagreement-centric adjudication stage, in which an independent adjudicator model examines the verified labels and justifications and determines a final label in accordance with the rubric. This structure mirrors established human annotation workflows in educational research, where initial coding is followed by self-checking and expert resolution of disagreements.

大模型教育数据多智能体标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。