arXiv:2410.13332cs.CL2024-10EMNLP被引 11

用多数据集联合微调提升小样本引文意图分类效果

Fine-Tuning Language Models on Multiple Datasets for Citation Intention Classification

  • 通过多任务学习融合主数据集与多个辅助数据集进行微调
  • 在小数据集上性能超越现有最佳模型7%~11%
  • 自适应调节辅助数据贡献,避免负迁移

引文意图分类(CIC)工具可识别引文目的(如背景、动机),帮助读者评估文献贡献。先前研究显示,预训练语言模型(PLM)如SciBERT在CIC基准上表现优异。这些模型通过大规模通用文本的自监督训练,经适度微调即可适配CIC任务。然而,微调时在小数据集上易过拟合。本文提出一种多任务学习(MTL)框架,联合主数据集与多个辅助CIC数据集进行微调,以利用额外监督信号。我们设计了一种数据驱动的任务关系学习(TRL)方法,动态控制辅助数据贡献,避免负迁移与繁琐超参数调优。在三个CIC数据集上实验表明,引入额外数据能提升主数据集的泛化性能。使用本框架微调的PLM在小数据集上比当前最优模型提升7%至11%,在大数据集上达到最佳水平。

原文摘要 · Abstract (English)

Citation intention Classification (CIC) tools classify citations by their intention (e.g., background, motivation) and assist readers in evaluating the contribution of scientific literature. Prior research has shown that pretrained language models (PLMs) such as SciBERT can achieve state-of-the-art performance on CIC benchmarks. PLMs are trained via self-supervision tasks on a large corpus of general text and can quickly adapt to CIC tasks via moderate fine-tuning on the corresponding dataset. Despite their advantages, PLMs can easily overfit small datasets during fine-tuning. In this paper, we propose a multi-task learning (MTL) framework that jointly fine-tunes PLMs on a dataset of primary interest together with multiple auxiliary CIC datasets to take advantage of additional supervision signals. We develop a data-driven task relation learning (TRL) method that controls the contribution of auxiliary datasets to avoid negative transfer and expensive hyper-parameter tuning. We conduct experiments on three CIC datasets and show that fine-tuning with additional datasets can improve the PLMs' generalization performance on the primary dataset. PLMs fine-tuned with our proposed framework outperform the current state-of-the-art models by 7% to 11% on small datasets while aligning with the best-performing model on a large dataset.

引文分析多任务学习小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。