arXiv:2410.02028cs.CL2024-10EMNLP被引 14

用大模型提升学术文档修改意图分类,提出新框架与数据集。

Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions

  • 对比生成式与编码式方法,系统评估大模型在分类任务中的表现。
  • 构建包含94,000条标注修改的1,780篇科学文档修订数据集。
  • 成果可复现,适合研究学术写作行为与模型分类能力的研究者。

分类是核心NLP任务,具有广泛应用潜力。尽管大语言模型(LLMs)在文本生成方面取得显著进展,其在分类任务中的潜力仍待深入探索。为填补这一空白,我们提出一个全面评估微调LLMs用于分类的框架,涵盖生成式与编码式方法。以挑战性且研究不足的编辑意图分类(EIC)为例,通过大量实验和多种训练策略及代表性LLMs的系统比较,获得关于其在EIC中应用的新见解。进一步验证这些发现的泛化能力于五个额外分类任务。为解决实证编辑分析的数据短缺问题,我们利用最优的EIC模型构建了Re3-Sci2.0,一个包含1,780篇科学文档修订、超过94,000条标注修改的大规模数据集,并通过人工评估确保质量。该数据集支持对学术写作中人类编辑行为的深入实证研究。相关实验框架、模型与数据均已公开。

原文摘要 · Abstract (English)

Classification is a core NLP task architecture with many potential applications. While large language models (LLMs) have brought substantial advancements in text generation, their potential for enhancing classification tasks remains underexplored. To address this gap, we propose a framework for thoroughly investigating fine-tuning LLMs for classification, including both generation- and encoding-based approaches. We instantiate this framework in edit intent classification (EIC), a challenging and underexplored classification task. Our extensive experiments and systematic comparisons with various training approaches and a representative selection of LLMs yield new insights into their application for EIC. We investigate the generalizability of these findings on five further classification tasks. To demonstrate the proposed methods and address the data shortage for empirical edit analysis, we use our best-performing EIC model to create Re3-Sci2.0, a new large-scale dataset of 1,780 scientific document revisions with over 94k labeled edits. The quality of the dataset is assessed through human evaluation. The new dataset enables an in-depth empirical study of human editing behavior in academic writing. We make our experimental framework, models and data publicly available.

大模型分类任务学术写作数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。