arXiv:2502.11916cs.CLcs.AI2025-02ACL被引 14

首个多粒度作文评分基准,评估大模型写作评估能力

EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models

  • 构建从词到篇章的多粒度评分框架,融合多模态上下文理解
  • 18个主流多模态大模型在篇章层面表现仍显著落后于人类
  • 适合研究教育评估、大模型可解释性与写作评价的学者

自动作文评分(AES)在教育评估中发挥重要作用,可提供可扩展且一致的写作任务评价。然而传统系统面临三大挑战:(1) 依赖手工特征,泛化能力有限;(2) 难以捕捉连贯性、论证等细粒度特质;(3) 无法处理多模态情境。在多模态大模型(MLLMs)时代,我们提出EssayJudge,首个用于评估多模态大模型作文评分能力的多粒度基准,覆盖词汇、句子和语篇层级特质。通过利用MLLM在特定特质评分与多模态上下文理解上的优势,EssayJudge实现无需人工特征工程的精准、上下文丰富评价,解决了长期存在的AES局限。对18个代表性MLLM的实验显示,其在人类评分对比下,尤其在语篇级特质上仍存在明显差距,凸显了基于MLLM的作文评分研究仍需进一步发展。

原文摘要 · Abstract (English)

Automated Essay Scoring (AES) plays a crucial role in educational assessment by providing scalable and consistent evaluations of writing tasks. However, traditional AES systems face three major challenges: (1) reliance on handcrafted features that limit generalizability, (2) difficulty in capturing fine-grained traits like coherence and argumentation, and (3) inability to handle multimodal contexts. In the era of Multimodal Large Language Models (MLLMs), we propose EssayJudge, the first multimodal benchmark to evaluate AES capabilities across lexical-, sentence-, and discourse-level traits. By leveraging MLLMs' strengths in trait-specific scoring and multimodal context understanding, EssayJudge aims to offer precise, context-rich evaluations without manual feature engineering, addressing longstanding AES limitations. Our experiments with 18 representative MLLMs reveal gaps in AES performance compared to human evaluation, particularly in discourse-level traits, highlighting the need for further advancements in MLLM-based AES research.

作文评分多模态大模型教育评估评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。