arXiv:2505.23818cs.CL2025-05被引 2

用生成式AI实现可解释的自动评分,支持多种题型和学科。

Ratas framework: A comprehensive genai-based approach to rubric-based marking of real-world textual exams

  • 基于生成式AI构建树状评分框架,适配各类评分标准。
  • 在真实项目课程数据上达到高可靠性和准确性。
  • 输出带理由的评分结果,适合教育评估与教学反馈。

自动化答案评分是教育技术中的关键挑战,具有提升评估效率、确保评分一致性并及时反馈学生的优势。然而,现有方法通常局限于特定考试格式,评分过程缺乏可解释性,且在跨学科和多样化评估类型中的实际应用能力不足。为此,我们提出RATAS(Rubric Automated Tree-based Answer Scoring)框架,利用前沿生成式AI模型实现基于评分量表的文本作答自动评分。RATAS支持广泛评分标准,具备学科无关的评估能力,并能生成结构化、可解释的评分理由。我们通过数学建模形式化自动评分任务,设计了可处理复杂真实考试结构的架构。为严格评估该方法,我们构建了一个源自真实项目课程的独特上下文数据集,涵盖多样响应格式与不同复杂度。实验结果表明,RATAS在自动化评分中展现出高可靠性与准确性,同时提供可解释反馈,增强学生与教师对评分透明度的信任。

原文摘要 · Abstract (English)

Automated answer grading is a critical challenge in educational technology, with the potential to streamline assessment processes, ensure grading consistency, and provide timely feedback to students. However, existing approaches are often constrained to specific exam formats, lack interpretability in score assignment, and struggle with real-world applicability across diverse subjects and assessment types. To address these limitations, we introduce RATAS (Rubric Automated Tree-based Answer Scoring), a novel framework that leverages state-of-the-art generative AI models for rubric-based grading of textual responses. RATAS is designed to support a wide range of grading rubrics, enable subject-agnostic evaluation, and generate structured, explainable rationales for assigned scores. We formalize the automatic grading task through a mathematical framework tailored to rubric-based assessment and present an architecture capable of handling complex, real-world exam structures. To rigorously evaluate our approach, we construct a unique, contextualized dataset derived from real-world project-based courses, encompassing diverse response formats and varying levels of complexity. Empirical results demonstrate that RATAS achieves high reliability and accuracy in automated grading while providing interpretable feedback that enhances transparency for both students and nstructors.

教育AI自动评分生成式AI可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。