arXiv:2601.20774cs.LG2026-01被引 1

数据再多也未必能提升多任务学习效果,因存在根本性适应瓶颈。

When More Data Doesn't Help: Limits of Adaptation in Multitask Learning

  • 分析多任务学习的统计极限,证明即使每任务数据无限多也无法突破适应瓶颈。
  • 揭示无分布信息时,单纯合并样本无法实现最优风险,与数据量无关。
  • 适用于研究多任务学习理论极限的学者,或需理解数据增益失效场景的工程师。

多任务学习及其相关框架在现代应用中取得了巨大成功。在多任务学习问题中,我们获得一组来自相关源任务的异构数据集,希望提升整体性能,超越单独解决每个任务所能达到的水平。先前工作(arXiv:2006.15785)表明,在缺乏分布信息的情况下,仅通过聚合样本的任何算法都无法保证最优风险,只要每任务的样本量有界。本文聚焦于理解多任务学习的统计极限,超越 arXiv:2006.15785 中的无免费午餐定理,建立了更强的适应性不可能性结果,该结果对任意大的每任务样本量均成立。这一改进传达出一个关键信息:多任务学习的困难无法通过增加每任务的数据量来克服。我们还讨论了未来可能感兴趣的最优适应性概念。

原文摘要 · Abstract (English)

Multitask learning and related frameworks have achieved tremendous success in modern applications. In multitask learning problem, we are given a set of heterogeneous datasets collected from related source tasks and hope to enhance the performance above what we could hope to achieve by solving each of them individually. The recent work of arXiv:2006.15785 has showed that, without access to distributional information, no algorithm based on aggregating samples alone can guarantee optimal risk as long as the sample size per task is bounded. In this paper, we focus on understanding the statistical limits of multitask learning. We go beyond the no-free-lunch theorem in arXiv:2006.15785 by establishing a stronger impossibility result of adaptation that holds for arbitrarily large sample size per task. This improvement conveys an important message that the hardness of multitask learning cannot be overcame by having abundant data per task. We also discuss the notion of optimal adaptivity that may be of future interests.

多任务学习统计极限适应性瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。