arXiv:2601.12618cs.CL2026-01被引 5

用大模型推理分歧做教育分析,让AI吵架变成有用数据。

Disagreement as Data: Reasoning Trace Analytics in Multi-Agent Systems

  • 用余弦相似度量化多智能体推理分歧,将争论转为分析信号。
  • 近万组对话编码中,模型分歧度与人类编码可靠性高度相关。
  • 适合教育研究者提升编码一致性,尤其人机协作场景下。

学习分析研究者常通过人工编码或访谈记录等定性数据理解学习过程。随着生成式AI兴起,全自动及人机协同分析流程逐渐成为可能,但方法论标准仍不健全。本研究提出:大型语言模型(LLM)智能体在多智能体系统中生成的推理轨迹,构成一种新型且丰富的过程数据,可增强定性编码的解释力。我们采用余弦相似度系统检测、量化并解读智能体间的分歧,将分歧重新定义为有意义的分析信号。通过对近10,000对智能体编码人类辅导对话片段的分析,发现模型推理语义相似度能稳健区分共识与分歧,并与人类编码可靠性显著相关。基于该指标的质性分析揭示了编码中的细微教学功能,也暴露了代码本可优化之处。结合定量相似度与定性审查,该方法有望提升编码过程中一致性建立的效率与准确性,尤其在人机协作中凸显解释性模糊。我们指出,推理轨迹分歧代表了一类重要的新分析信号,有助于推动教育研究的方法严谨性与解释深度。

原文摘要 · Abstract (English)

Learning analytics researchers often analyze qualitative student data such as coded annotations or interview transcripts to understand learning processes. With the rise of generative AI, fully automated and human-AI workflows have emerged as promising methods for analysis. However, methodological standards to guide such workflows remain limited. In this study, we propose that reasoning traces generated by large language model (LLM) agents, especially within multi-agent systems, constitute a novel and rich form of process data to enhance interpretive practices in qualitative coding. We apply cosine similarity to LLM reasoning traces to systematically detect, quantify, and interpret disagreements among agents, reframing disagreement as a meaningful analytic signal. Analyzing nearly 10,000 instances of agent pairs coding human tutoring dialog segments, we show that LLM agents' semantic reasoning similarity robustly differentiates consensus from disagreement and correlates with human coding reliability. Qualitative analysis guided by this metric reveals nuanced instructional sub-functions within codes and opportunities for conceptual codebook refinement. By integrating quantitative similarity metrics with qualitative review, our method has the potential to improve and accelerate establishing inter-rater reliability during coding by surfacing interpretive ambiguity, especially when LLMs collaborate with humans. We discuss how reasoning-trace disagreements represent a valuable new class of analytic signals advancing methodological rigor and interpretive depth in educational research.

多智能体推理分析教育研究大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。