arXiv:2511.20547cs.CL2025-11被引 2

构建学生对话标注数据集,用大模型自动识别知识建构与任务完成话语。

From Words to Wisdom: Discourse Annotation and Baseline Models for Student Dialogue Understanding

  • 基于真实课堂对话构建标注数据集,聚焦知识建构与任务完成两类话语。
  • 使用GPT-3.5和Llama-3.1进行基线预测,性能未达理想水平。
  • 为教育对话分析提供可扩展的自动化工具,适合教育技术研究者使用。

识别学生对话中的语篇特征对教育研究者理解知识建构而非仅完成任务的教学变量至关重要。人工分析此类对话耗时费力,限制了研究规模。利用自然语言处理技术可实现语篇特征的自动检测,为教育研究提供可扩展的数据驱动洞察。然而,现有关注对话语篇的NLP研究极少涉及教育数据。本文通过引入一个包含知识建构与任务生产话语的学生对话标注数据集,并基于预训练大语言模型GPT-3.5和Llama-3.1建立基线模型,实现对话中每一轮发言的语篇属性自动预测。实验表明,当前先进模型在该任务上表现不佳,提示未来研究空间。

原文摘要 · Abstract (English)

Identifying discourse features in student conversations is quite important for educational researchers to recognize the curricular and pedagogical variables that cause students to engage in constructing knowledge rather than merely completing tasks. The manual analysis of student conversations to identify these discourse features is time-consuming and labor-intensive, which limits the scale and scope of studies. Leveraging natural language processing (NLP) techniques can facilitate the automatic detection of these discourse features, offering educational researchers scalable and data-driven insights. However, existing studies in NLP that focus on discourse in dialogue rarely address educational data. In this work, we address this gap by introducing an annotated educational dialogue dataset of student conversations featuring knowledge construction and task production discourse. We also establish baseline models for automatically predicting these discourse properties for each turn of talk within conversations, using pre-trained large language models GPT-3.5 and Llama-3.1. Experimental results indicate that these state-of-the-art models perform suboptimally on this task, indicating the potential for future research.

对话分析教育AI大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。