arXiv:2604.09622cs.CYcs.AI2026-04被引 1

为AI生成的试题提供可解释与认证框架,提升教育评估可信度。

Explainability and Certification of AI-Generated Educational Assessments

  • 结合自解释、归因分析与事后验证,生成可解释的认知对齐证据。
  • 500道计算机科学试题测试显示透明度提升,教师工作量降低。
  • 提出交通灯认证流程,适合教育机构和认证体系使用。

生成式人工智能在教育评估中的快速应用带来了大规模题库生成、个性化反馈和高效形成性评价的新机遇。然而,尽管在分类体系对齐和自动出题方面取得进展,缺乏透明、可解释且可认证的机制仍限制了机构及认证层面的采纳。本文提出一个综合性的可解释性与认证框架,融合自解释、基于归因的分析以及事后验证,生成基于布卢姆分类法和SOLO分类法的认知对齐证据。引入结构化的认证元数据模式,记录来源、对齐预测、评审动作与伦理指标,实现符合新兴治理要求的可审计文档。通过交通灯认证流程,将题目分为可自动认证、需人工审查或拒绝三类。对500道AI生成的计算机科学试题进行概念验证研究,结果表明该框架具备可行性,显著提升透明度、减少教师负担并增强可审计性。文章最后讨论了伦理影响、政策考量及未来研究方向,强调可解释性与认证是构建可信、可认证的AI评估系统的核心要素。

原文摘要 · Abstract (English)

The rapid adoption of generative artificial intelligence (AI) in educational assessment has created new opportunities for scalable item creation, personalized feedback, and efficient formative evaluation. However, despite advances in taxonomy alignment and automated question generation, the absence of transparent, explainable, and certifiable mechanisms limits institutional and accreditation-level acceptance. This chapter proposes a comprehensive framework for explainability and certification of AI-generated assessment items, combining self-rationalization, attribution-based analysis, and post-hoc verification to produce interpretable cognitive-alignment evidence grounded in Bloom's and SOLO taxonomies. A structured certification metadata schema is introduced to capture provenance, alignment predictions, reviewer actions, and ethical indicators, enabling audit-ready documentation consistent with emerging governance requirements. A traffic-light certification workflow operationalizes these signals by distinguishing auto-certifiable items from those requiring human review or rejection. A proof-of-concept study on 500 AI-generated computer science questions demonstrates the framework's feasibility, showing improved transparency, reduced instructor workload, and enhanced auditability. The chapter concludes by outlining ethical implications, policy considerations, and directions for future research, positioning explainability and certification as essential components of trustworthy, accreditation-ready AI assessment systems.

AI评估可解释性认证框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。