用大模型统一评主观题,像真人一样判断答案质量。
Towards Human-Like Grading: A Unified LLM-Enhanced Framework for Subjective Question Evaluation
- 融合大模型推理与生成能力,分四模块综合评估答案
- 在多个数据集上优于传统和现有大模型基线方法
- 已落地电商企业培训认证考试,支持多类型主观题
主观题自动评分仍是考试评估中的重大挑战,因题目形式多样、学生作答开放。现有研究多聚焦特定题型,缺乏对包含多种题型的综合性考试的支持。本文提出一种统一的大语言模型(LLM)增强自动评分框架,可对跨领域各类主观题提供类人评价。该框架集成四个互补模块:基础文本匹配模块提供内容相似性评估;利用LLM能力实现:(1)比对学生答案与参考答案的关键知识点;(2)从学生答案生成伪问题以评估相关性;(3)模拟人工评分,识别内容相关及非内容层面的优缺点。在通用与领域特定数据集上的大量实验表明,本框架在多个评分指标上持续优于传统及基于LLM的基线方法。此外,该系统已在一家大型电商平台的真实培训与认证考试中成功部署。
原文摘要 · Abstract (English)
Automatic grading of subjective questions remains a significant challenge in examination assessment due to the diversity in question formats and the open-ended nature of student responses. Existing works primarily focus on a specific type of subjective question and lack the generality to support comprehensive exams that contain diverse question types. In this paper, we propose a unified Large Language Model (LLM)-enhanced auto-grading framework that provides human-like evaluation for all types of subjective questions across various domains. Our framework integrates four complementary modules to holistically evaluate student answers. In addition to a basic text matching module that provides a foundational assessment of content similarity, we leverage the powerful reasoning and generative capabilities of LLMs to: (1) compare key knowledge points extracted from both student and reference answers, (2) generate a pseudo-question from the student answer to assess its relevance to the original question, and (3) simulate human evaluation by identifying content-related and non-content strengths and weaknesses. Extensive experiments on both general-purpose and domain-specific datasets show that our framework consistently outperforms traditional and LLM-based baselines across multiple grading metrics. Moreover, the proposed system has been successfully deployed in real-world training and certification exams at a major e-commerce enterprise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。