首个针对中文多文体作文的评测基准,覆盖4类文体。
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
- 构建4类中文作文(议论文/记叙文/描写文/说明文)评测集
- 使用728个真实题目,分开放与受限两类场景
- 设计细粒度评分框架,支持跨文体客观评估
中文作文写作及其评估在教育领域至关重要,但大语言模型(LLMs)在此领域的表现仍缺乏深入研究。现有评测基准多依赖粗粒度文本质量指标,忽视了中文作文在结构和修辞上的复杂性,尤其在不同文体间的差异。为此,我们提出 enchName,一个专为中文作文设计的多文体评测基准,涵盖议论文、记叙文、描写文和说明文四类主要文体。我们收集并精炼了728个真实作文题目,确保真实性,并将其细致划分为开放式与受限式两类,以覆盖多样化的写作情境。为实现可靠评估,我们开发了一套细粒度、文体特异的评分框架,采用分层聚合方式生成综合得分。通过全面的人工一致性研究验证了评估协议的有效性。最后,我们对15个大规模语言模型进行了评测,分析其在不同文体和指令类型下的优劣势。本工作旨在推动基于LLM的中文作文评估发展,并激励未来在教育场景下提升作文生成能力的研究。
原文摘要 · Abstract (English)
Chinese essay writing and its evaluation are critical in educational contexts, yet the capabilities of Large Language Models (LLMs) in this domain remain largely underexplored. Existing benchmarks often rely on coarse-grained text quality metrics, largely overlooking the structural and rhetorical complexities of Chinese essays, particularly across diverse genres. To address this gap, we propose \benchName, a multi-genre benchmark specifically designed for Chinese essay writing across four major genres: Argumentative, Narrative, Descriptive, and Expository. We curate and refine a total of 728 real-world prompts to ensure authenticity and meticulously categorize them into the \textit{Open-Ended} and \textit{Constrained} sets to capture diverse writing scenarios. To reliably evaluate generated essays, we develop a fine-grained, genre-specific scoring framework that hierarchically aggregates scores. We further validate our evaluation protocol through a comprehensive human agreement study. Finally, we benchmark 15 large-sized LLMs, analyzing their strengths and limitations across genres and instruction types. With \benchName, we aim to advance LLM-based Chinese essay evaluation and inspire future research on improving essay generation in educational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。