arXiv:2409.20042cs.CLcs.AI2024-09被引 18

零样本生成答案评分与反馈,省去调优还更准

Beyond Scores: A Modular RAG-Based System for Automatic Short Answer Scoring with Feedback

  • 用模块化检索增强生成,零样本/少样本下自动评分
  • 在未见过题目上评分准确率提升9%,超越微调方法
  • 适合教育场景快速部署,无需复杂提示工程

自动短答案评分(ASAS)可减轻教师批改负担,但现有带反馈的ASAS方法常依赖小规模数据微调语言模型,成本高且泛化能力差。近期基于大模型的方法虽减少微调,却多依赖提示工程,难以生成详实反馈或有效评估反馈质量。本文提出一种基于模块化检索增强生成的ASAS-F系统,在严格零样本和少样本条件下实现答案评分与反馈生成。通过自动提示生成框架,系统可适配多种教育任务而无需大量提示工程。实验表明,在未见题目上评分准确率相比微调方法提升9%,提供了一种可扩展、低成本的解决方案。

原文摘要 · Abstract (English)

Automatic short answer scoring (ASAS) helps reduce the grading burden on educators but often lacks detailed, explainable feedback. Existing methods in ASAS with feedback (ASAS-F) rely on fine-tuning language models with limited datasets, which is resource-intensive and struggles to generalize across contexts. Recent approaches using large language models (LLMs) have focused on scoring without extensive fine-tuning. However, they often rely heavily on prompt engineering and either fail to generate elaborated feedback or do not adequately evaluate it. In this paper, we propose a modular retrieval augmented generation based ASAS-F system that scores answers and generates feedback in strict zero-shot and few-shot learning scenarios. We design our system to be adaptable to various educational tasks without extensive prompt engineering using an automatic prompt generation framework. Results show an improvement in scoring accuracy by 9\% on unseen questions compared to fine-tuning, offering a scalable and cost-effective solution.

自动评分RAG教育AI零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。