arXiv:2602.15034cs.CLcs.AI2026-02

首个面向教育研究全流程的细粒度评估基准,助力LLM学术写作能力精准诊断。

EduResearchBench: A Hierarchical Atomic Task Decomposition Benchmark for Full-Lifecycle Educational Research

  • 将教育研究分解为24个原子任务,构建分层评估框架
  • 用11K高质量指令对训练出的EduWrite模型超越72B通用模型
  • 适合需要精细化评估学术写作能力的研究者使用

尽管大语言模型正在重塑社会科学中的人工智能应用,但对其学术写作能力的严格评估仍是重大挑战。现有基准多聚焦单次生成任务,缺乏对复杂科研流程的细粒度评价。为此,我们提出EduResearchBench,首个专用于教育学术写作的综合性评估平台。该平台基于分层原子任务分解(HATD)框架,将完整研究流程拆分为六个专业模块(如定量分析、质性研究、政策研究),涵盖24个细粒度原子任务。这一分类体系支持自动化评估流程,克服了整体评分掩盖具体短板的问题,提供可诊断的缺陷反馈。针对学术写作的高认知负荷,我们设计渐进式课程学习策略,从基础技能逐步过渡到复杂方法推理与论证。基于5.5万原始学术样本,我们构建了1.1万条高质量指令对,用于训练专用教育学术写作模型EduWrite。实验表明,EduWrite(30B)在多个核心指标上显著优于更大规模的通用模型(72B),证明在垂直领域中,数据质量密度与分阶段训练课程比参数规模更具决定性作用。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) are reshaping the paradigm of AI for Social Science (AI4SS), rigorously evaluating their capabilities in scholarly writing remains a major challenge. Existing benchmarks largely emphasize single-shot, monolithic generation and thus lack the fine-grained assessments required to reflect complex academic research workflows. To fill this gap, we introduce EduResearchBench, the first comprehensive evaluation platform dedicated to educational academic writing. EduResearchBench is built upon our Hierarchical Atomic Task Decomposition (HATD) framework, which decomposes an end-to-end research workflow into six specialized research modules (e.g., Quantitative Analysis, Qualitative Research, and Policy Research) spanning 24 fine-grained atomic tasks. This taxonomy enables an automated evaluation pipeline that mitigates a key limitation of holistic scoring, where aggregate scores often obscure specific capability bottlenecks, and instead provides fine-grained, diagnostic feedback on concrete deficiencies. Moreover, recognizing the high cognitive load inherent in scholarly writing, we propose a curriculum learning strategy that progressively builds competence from foundational skills to complex methodological reasoning and argumentation. Leveraging 55K raw academic samples, we curate 11K high-quality instruction pairs to train EduWrite, a specialized educational scholarly writing model. Experiments show that EduWrite (30B) substantially outperforms larger general-purpose models (72B) on multiple core metrics, demonstrating that in vertical domains, data quality density and hierarchically staged training curricula are more decisive than parameter scale.

教育AILLM评估学术写作细粒度评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。