分析迁移学习中自适应数据选择的可复现性问题,揭示性能与稳定性权衡机制。
Sensitivity of Stability: Theoretical & Empirical Analysis of Replicability for Adaptive Data Selection in Transfer Learning
- 提出选择敏感度Δ_Q,量化数据选择策略对训练扰动的响应程度。
- 高自适应策略性能优但复现失败率超30%,低自适应策略失败率低于7%。
- 预训练源域数据可降低30%失败率,兼顾性能与可复现性。
迁移学习的广泛应用已革新机器学习,实现预训练模型向新领域的高效适配。然而,采用自适应数据选择策略时,其结果可靠性仍不明确。本文提出理论与实证结合的分析框架,建立适应性有效性与结果一致性之间的根本权衡。核心贡献是形式化选择敏感度(Δ_Q),衡量自适应选择策略对训练数据扰动的响应程度。证明复现失败概率——两次独立训练产生性能差异超过阈值的概率——随Δ_Q平方增长,随样本量指数下降。在MultiNLI语料库上,六种自适应策略(从均匀采样到基于梯度的选择)的实验验证了该理论关系。结果显示,梯度和课程学习等高度自适应策略虽性能优异,但复现失败率显著升高;而低自适应方法失败率稳定在7%以下。关键发现:源域预训练可将失败率降低最高30%,同时保持性能优势。研究为实践者提供了性能-可复现性权衡的指导原则,并呼吁设计更具备可复现意识的现代迁移学习系统。
原文摘要 · Abstract (English)
The widespread adoption of transfer learning has revolutionized machine learning by enabling efficient adaptation of pre-trained models to new domains. However, the reliability of these adaptations remains poorly understood, particularly when using adaptive data selection strategies that dynamically prioritize training examples. We present a comprehensive theoretical and empirical analysis of replicability in transfer learning, introducing a mathematical framework that quantifies the fundamental trade-off between adaptation effectiveness and result consistency. Our key contribution is the formalization of selection sensitivity ($Δ_Q$), a measure that captures how adaptive selection strategies respond to perturbations in training data. We prove that replicability failure probability: the likelihood that two independent training runs produce models differing in performance by more than a threshold, increases quadratically with selection sensitivity while decreasing exponentially with sample size. Through extensive experiments on the MultiNLI corpus using six adaptive selection strategies - ranging from uniform sampling to gradient-based selection - we demonstrate that this theoretical relationship holds precisely in practice. Our results reveal that highly adaptive strategies like gradient-based and curriculum learning achieve superior task performance but suffer from high replicability failure rates, while less adaptive approaches maintain failure rates below 7%. Crucially, we show that source domain pretraining provides a powerful mitigation mechanism, reducing failure rates by up to 30% while preserving performance gains. These findings establish principled guidelines for practitioners to navigate the performance-replicability trade-off and highlight the need for replicability-aware design in modern transfer learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。