arXiv:2604.23449cs.AIcs.HC2026-04中稿 · the 27th Internati…

用AI实时分组,让不同水平学生在科学讨论中更平等参与。

ArguAgent: AI-Supported Real-Time Grouping for Productive Argumentation in STEM Classrooms

  • 基于立场和论证能力动态分组,确保多样性与能力均衡。
  • 模拟测试显示95.4%分组符合设计标准,是随机分组的3.2倍效率。
  • 适合需要促进课堂有效讨论的中小学科学教师使用。

论证是科学教育的核心实践,但其成效取决于参与者及其互动方式。高成就学生常主导讨论,低成就学生则可能沉默或被动接受。若能根据学生观点和论证能力实时分组,可促进包容性、基于证据的对话。然而教师难以在教学中实时获取可靠的学生立场与论证质量信息。本文提出生成式AI系统ArguAgent,通过双阶段评估流程:先对论点按0-4分量表评分,再基于语义分析聚类立场。该评分模块经200条专家标注验证,与人工共识一致性达Krippendorff's α = 0.817。对比GPT-4o-mini、GPT-5.1、GPT-5.2三模型,经人类分歧分析优化提示词使评分提升89%(QWK: 0.531→0.686),模型升级贡献11%(QWK: 0.686→0.708)。在100个班级的模拟测试中,分组算法实现95.4%的组别满足立场异质性与能力差异≤±1级的双重标准,较随机分组提升3.2倍。结果表明,ArguAgent可支持实时、有理论依据的分组,推动课堂教学中的高效论证。

原文摘要 · Abstract (English)

Argumentation is a core practice in STEM education, but its productivity depends on who participates and how they interact. Higher-achieving students often dominate the talk and decision-making, while lower-achieving peers may disengage, defer, or comply without contributing substantive reasoning. Forming groups strategically based on students' stances and argumentation skills could help foster inclusive, evidence-based discourse. In practice, however, teachers are constrained in implementing this grouping strategy because it requires real-time insight into students' positions and the quality of their argumentation, information that is difficult to assess reliably and at scale during instruction. We present a generative AI-powered system, ArguAgent, that creates groups optimizing for stance heterogeneity while constraining argumentation quality differences to +/-1 level on a validated learning progression. ArguAgent uses a two-component assessment pipeline: first scoring student arguments on a 0-4 rubric, then clustering positions via semantic analysis. We validated the scoring component against human expert consensus (Krippendorff's ααα = 0.817) using 200 expert-generated scores. Testing three OpenAI models (GPT-4o-mini, GPT-5.1, GPT-5.2) with identical calibrated prompts, we found that systematic prompt engineering informed by human disagreement analysis contributed 89% of scoring improvement (QWK: 0.531 to 0.686), while model upgrades contributed an additional 11% (QWK: 0.686 to 0.708). Simulation testing across 100 classes demonstrated that the grouping algorithm achieves 95.4% of groups that meet both design criteria, a 3.2x improvement over random assignment. These results suggest ArguAgent can enable real-time, theoretically grounded grouping that promotes productive STEM argumentation in classrooms.

AI教育课堂分组论证能力STEM教学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。