arXiv:2507.15236cs.CLcs.LG2025-07

通过兴趣子集分析多场景训练中模型学习行为变化。

SOI Matters: Analyzing Multi-Setting Training Dynamics in Pretrained Language Models via Subsets of Interest

  • 提出SOI框架,识别六类学习模式,追踪例子在训练中的状态演变。
  • 多源学习使分布外性能提升7%,多任务学习效果取决于任务相似性。
  • 基于SOI的两阶段微调可进一步提升性能,适合优化多场景模型。

本研究探讨多任务、多语言、多源学习对预训练语言模型鲁棒性与性能的影响。为增强分析,提出兴趣子集(SOI)分类框架,识别训练过程中六种典型学习行为模式,包括易遗忘样本、未学习样本和始终正确样本。通过SOI转移热图与数据集制图可视化,分析从单设置到多设置配置时样本状态的变化。我们在三个平行对比实验中展开:使用英文任务(蕴含、改写、情感)比较多任务与单任务学习;使用情感分析数据集比较多源与单源学习;使用法语、英语、波斯语意图分类任务比较多语言与单语言学习。结果表明,多源学习可使分布外性能提升最高达7%,多任务学习效果混合,但在任务相似组合中表现显著。我们进一步提出两阶段微调方法,第二阶段利用SOI子集选择实现额外性能提升。研究揭示了训练动态新机制,并为优化多设置语言模型性能提供实用路径。

原文摘要 · Abstract (English)

This work investigates the impact of multi-task, multi-lingual, and multi-source learning approaches on the robustness and performance of pretrained language models. To enhance this analysis, we introduce Subsets of Interest (SOI), a novel categorization framework that identifies six distinct learning behavior patterns during training, including forgettable examples, unlearned examples, and always correct examples. Through SOI transition heatmaps and dataset cartography visualization, we analyze how examples shift between these categories when transitioning from single-setting to multi-setting configurations. We perform comprehensive experiments across three parallel comparisons: multi-task vs. single-task learning using English tasks (entailment, paraphrase, sentiment), multi-source vs. single-source learning using sentiment analysis datasets, and multi-lingual vs. single-lingual learning using intent classification in French, English, and Persian. Our results demonstrate that multi-source learning consistently improves out-of-distribution performance by up to 7%, while multi-task learning shows mixed results with notable gains in similar task combinations. We further introduce a two-stage fine-tuning approach where the second stage leverages SOI-based subset selection to achieve additional performance improvements. These findings provide new insights into training dynamics and offer practical approaches for optimizing multi-setting language model performance.

预训练模型多任务学习子集分析训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。