提出教育AI评估新框架,关注育人价值而非仅技术指标。
Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
- 构建十维可衡量的TEACH-AI评估框架,融合教育学与社会技术视角。
- 强调学习者主体性与伦理,推动人机协同教育生态建设。
- 适合教育AI设计者、研究者与政策制定者参考使用。
随着生成式人工智能持续重塑教育,现有评估多聚焦于准确性或任务效率等技术指标,忽视了学习者身份、自主性、情境化学习过程及伦理考量。本文提出TEACH-AI(可信且有效的教育课堂启发式)框架,一个领域无关、以教育学为基础、利益相关者对齐的评估体系,包含可测量指标与实用工具包,支持生成式AI在教育场景中的设计、开发与评估。该框架基于广泛文献综述与整合,提供十组件评估体系与检查清单,为可扩展、价值观一致的教育AI评估奠定基础。TEACH-AI从社会技术、教育理论与应用多维度重构‘评估’内涵,促进教育与AI领域设计者、开发者、研究者与政策制定者的协同参与。本工作呼吁学界重新思考教育中‘有效AI’的定义,倡导以共创、包容和长期人文、社会与教育影响为导向的模型评估方法。
原文摘要 · Abstract (English)
As generative artificial intelligence (AI) continues to transform education, most existing AI evaluations rely primarily on technical performance metrics such as accuracy or task efficiency while overlooking human identity, learner agency, contextual learning processes, and ethical considerations. In this paper, we present TEACH-AI (Trustworthy and Effective AI Classroom Heuristics), a domain-independent, pedagogically grounded, and stakeholder-aligned framework with measurable indicators and a practical toolkit for guiding the design, development, and evaluation of generative AI systems in educational contexts. Built on an extensive literature review and synthesis, the ten-component assessment framework and toolkit checklist provide a foundation for scalable, value-aligned AI evaluation in education. TEACH-AI rethinks "evaluation" through sociotechnical, educational, theoretical, and applied lenses, engaging designers, developers, researchers, and policymakers across AI and education. Our work invites the community to reconsider what constructs "effective" AI in education and to design model evaluation approaches that promote co-creation, inclusivity, and long-term human, social, and educational impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。