arXiv:2503.00867cs.CL2025-03被引 4

DUAL通过融合不确定性与多样性,提升文本摘要的主动学习效率。

DUAL: Diversity and Uncertainty Active Learning for Text Summarization

  • 结合不确定性和多样性选择样本,兼顾代表性与挑战性。
  • 在多个模型和数据集上表现优于随机采样及现有方法。
  • 适合资源有限时构建高质量摘要数据集的研究者使用。

随着大语言模型的发展,神经文本摘要近年来取得显著进展。然而,即使是顶尖模型仍严重依赖高质量人工标注数据进行训练与评估。主动学习常被用于高效收集此类数据,尤其在标注资源稀缺时。现有主动学习方法通常只侧重不确定性或多样性,但在摘要任务中效果有限,常不如随机采样。本文提出多样性与不确定性主动学习(DUAL),一种新算法,通过迭代选择既代表数据分布又对当前模型具挑战性的样本。DUAL缓解了基于不确定性的方法易选噪声样本的问题,以及基于多样性的方法探索范围有限的缺陷。在不同摘要模型与基准数据集上的大量实验表明,DUAL始终达到或超越最佳策略表现。通过可视化与量化指标,我们深入分析了不同主动学习策略在摘要任务中的有效性和鲁棒性,揭示其性能不一致的原因。最终,DUAL在多样性与鲁棒性之间取得了良好平衡。

原文摘要 · Abstract (English)

With the rise of large language models, neural text summarization has advanced significantly in recent years. However, even state-of-the-art models continue to rely heavily on high-quality human-annotated data for training and evaluation. Active learning is frequently used as an effective way to collect such datasets, especially when annotation resources are scarce. Active learning methods typically prioritize either uncertainty or diversity but have shown limited effectiveness in summarization, often being outperformed by random sampling. We present Diversity and Uncertainty Active Learning (DUAL), a novel algorithm that combines uncertainty and diversity to iteratively select and annotate samples that are both representative of the data distribution and challenging for the current model. DUAL addresses the selection of noisy samples in uncertainty-based methods and the limited exploration scope of diversity-based methods. Through extensive experiments with different summarization models and benchmark datasets, we demonstrate that DUAL consistently matches or outperforms the best performing strategies. Using visualizations and quantitative metrics, we provide valuable insights into the effectiveness and robustness of different active learning strategies, in an attempt to understand why these strategies haven't performed consistently in text summarization. Finally, we show that DUAL strikes a good balance between diversity and robustness.

主动学习文本摘要多样性不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。