arXiv:2602.18451cs.CYcs.AI2026-02

用多智能体系统自动生成符合NGSS标准的科学测评题

Developing a Multi-Agent System to Generate Next Generation Science Assessments with Evidence-Centered Design

  • 将证据中心设计框架融入多智能体系统,自动完成测评题生成流程
  • 生成题目与人工题在标准契合度和认知要求上基本相当
  • AI题更包容但清晰度、简洁性不足,需人类专家补足

当前科学教育改革如下一代科学标准(NGSS)要求评估学生运用科学知识解决问题和设计解决方案的能力。为考察这种高阶能力,教育者需要表现性评估,但其开发难度大。一种广泛应用的解决方案是证据中心设计(ECD),强调学习者、证据与任务之间的关联模型。尽管ECD能保障评估有效性,但其实现需跨领域专业知识(如内容与评估),成本高且耗时。为此,本研究提出将ECD框架整合至多智能体系统(MAS),实现NGSS对齐测评题的自动化生成。该系统融合多个具备不同专长的大语言模型,可自动执行传统上由人类专家完成的复杂多阶段题目生成工作。我们评估了AI生成题目的质量,并与人工开发题目在评估设计多个维度进行对比。结果显示,AI生成题目在与NGSS三维标准的契合度和认知需求方面总体与人工题相当;但存在差异:AI题在包容性上表现更优,而在清晰度、简洁性和多模态设计上仍有局限。无论是AI还是人工题,在证据可收集性和学生兴趣契合度方面均存在不足。这些发现表明,将ECD融入MAS可支持可扩展且符合标准的评估设计,但人类专家的作用仍不可替代。

原文摘要 · Abstract (English)

Contemporary science education reforms such as the Next Generation Science Standards (NGSS) demand assessments to understand students' ability to use science knowledge to solve problems and design solutions. To elicit such higher-order ability, educators need performance-based assessments, which are challenging to develop. One solution that has been broadly adopted is Evidence-Centered Design (ECD), which emphasizes interconnected models of the learner, evidence, and tasks. Although ECD provides a framework to safeguard assessment validity, its implementation requires diverse expertise (e.g., content and assessment), which is both costly and labor-intensive. To address this challenge, this study proposed integrating the ECD framework into Multi-Agent Systems (MAS) to generate NGSS-aligned assessment items automatically. This integrated MAS system ensembles multiple large language models with varying expertise, enabling the automation of complex, multi-stage item generation workflows traditionally performed by human experts. We examined the quality of AI-generated NGSS-aligned items and compared them with human-developed items across multiple dimensions of assessment design. Results showed that AI-generated items have overall comparable quality to human-developed items in terms of alignment with NGSS three-dimensional standards and cognitive demands. Divergent patterns also emerged: AI-generated items demonstrated a distinct strength in inclusivity, while also exhibiting limitations in clarity, conciseness, and multimodal design. AI- and human-developed items both showed weaknesses in evidence collectability and student interest alignment. These findings suggest that integrating ECD into MAS can support scalable and standards-aligned assessment design, while human expertise remains essential.

教育评估多智能体NGSSAI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。