arXiv:2410.09507cs.CL2024-10EMNLP被引 2

AERA Chat 用多模型生成答案评分与解释,让教师能直观评估学生作答逻辑。

AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment

  • 多大模型并行评分并生成解释,可视化关键答题点与理由
  • 支持教师标注与研究者对比不同模型的解释质量
  • 解决解释文本噪声问题,提升评分可信度与教学可用性

自动化学生作答评分中的可解释性对建立教育者信任和提升系统可用性至关重要。然而,由于标注数据稀缺和人工验证成本高昂,高质量解释的生成仍具挑战性,目前严重依赖大语言模型(LLMs)生成的解释,但这些解释常存在噪声且不可靠。为此,我们提出 AERA Chat,一个面向自动化可解释学生作答评估的交互式可视化平台。该平台利用多个 LLM 并行对学生的答案进行评分并生成解释性理由,提供创新的可视化功能,突出关键答题内容与解释依据。平台还集成直观的标注与评估工具,支持教育者完成标记任务,也便于研究者从不同模型中评估解释质量。我们在多个数据集上对多种解释生成方法进行了评估,验证了平台在促进稳健解释评估与对比分析方面的有效性。

原文摘要 · Abstract (English)

Explainability in automated student answer scoring systems is critical for building trust and enhancing usability among educators. Yet, generating high-quality assessment rationales remains challenging due to the scarcity of annotated data and the prohibitive cost of manual verification, prompting heavy reliance on rationales produced by large language models (LLMs), which are often noisy and unreliable. To address these limitations, we present AERA Chat, an interactive visualization platform designed for automated explainable student answer assessment. AERA Chat leverages multiple LLMs to concurrently score student answers and generate explanatory rationales, offering innovative visualization features that highlight critical answer components and rationale justifications. The platform also incorporates intuitive annotation and evaluation tools, supporting educators in marking tasks and researchers in evaluating rationale quality from different models. We demonstrate the effectiveness of our platform through evaluations of multiple rationale-generation methods on several datasets, showcasing its capability for facilitating robust rationale evaluation and comparative analysis.

可解释性教育AILLM应用交互平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。