arXiv:2511.09785cs.AI2025-11被引 7

用自检和互检提升大模型标注学习对话的准确性。

AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics

  • 让大模型自己检查或互相审核标注结果,提升质量。
  • 自检使标注一致性提升近一倍,复杂教学行为改善最明显。
  • 提出可复现的标注框架和标准化报告格式,适合教育数据研究者。

大语言模型在学习互动标注中应用日益广泛,但可靠性问题限制其使用。本研究测试了以验证为导向的编排策略——即提示模型自我检查(自检)或相互审计(互检)——是否能提升辅导对话的定性编码质量。基于30个一对一数学辅导会话的转录文本,比较了三种主流大模型(GPT、Claude、Gemini)在三种条件下的表现:无验证、自检与互检。所有编排配置均与盲评、以分歧为核心的真人标注进行对比,采用Cohen's kappa评估。结果显示,整体编排使kappa值提升58%;自检将一致性几乎翻倍,对复杂教学行为增益最大;互检平均提升37%,但效果依赖于验证者与标注者配对及具体概念,部分组合优于自检,部分反而降低一致性,反映验证者严格度差异。本文贡献包括:(1) 可灵活实现控制、自检与互检的编排框架;(2) 在真实辅导数据上对前沿大模型的实证比较,采用盲评人类“金标准”;(3) 提出简洁标注记法如Gemini(GPT),标准化报告并明确方向性效应,便于复现。结果表明,验证机制是实现可靠、可扩展的教育数据分析中大模型辅助标注的关键设计杠杆。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used to annotate learning interactions, yet concerns about reliability limit their utility. We test whether verification-oriented orchestration-prompting models to check their own labels (self-verification) or audit one another (cross-verification)-improves qualitative coding of tutoring discourse. Using transcripts from 30 one-to-one math sessions, we compare three production LLMs (GPT, Claude, Gemini) under three conditions: unverified annotation, self-verification, and cross-verification across all orchestration configurations. Outputs are benchmarked against a blinded, disagreement-focused human adjudication using Cohen's kappa. Overall, orchestration yields a 58 percent improvement in kappa. Self-verification nearly doubles agreement relative to unverified baselines, with the largest gains for challenging tutor moves. Cross-verification achieves a 37 percent improvement on average, with pair- and construct-dependent effects: some verifier-annotator pairs exceed self-verification, while others reduce alignment, reflecting differences in verifier strictness. We contribute: (1) a flexible orchestration framework instantiating control, self-, and cross-verification; (2) an empirical comparison across frontier LLMs on authentic tutoring data with blinded human "gold" labels; and (3) a concise notation, verifier(annotator) (e.g., Gemini(GPT) or Claude(Claude)), to standardize reporting and make directional effects explicit for replication. Results position verification as a principled design lever for reliable, scalable LLM-assisted annotation in Learning Analytics.

大模型教育数据标注质量自检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。