提出新框架提升多任务语言模型的鲁棒性与效率
An Information-theoretic Multi-task Representation Learning Framework for Natural Language Understanding
- 通过共享信息最大化学习通用表示,避免压缩导致的信息不足
- 利用任务特异性信息最小化去除冗余特征,提升预测准确性
- 在数据少和噪声多场景下表现优异,适合资源受限的NLP应用
本文提出一种基于信息论的多任务表示学习框架(InfoMTL),旨在为所有任务提取抗噪且充分的共享表示。该框架确保共享表示对所有任务都足够充分,并缓解冗余特征的负面影响,从而增强预训练语言模型在多任务设置下的语言理解能力。首先,提出共享信息最大化原则,以学习更充分的共享表示,避免多任务范式中因表示压缩引发的信息不足问题。其次,设计任务特异性信息最小化原则,以消除输入中潜在的冗余信息,保留与目标任务相关的必要信息,实现高效多任务预测。在六个分类基准上的实验表明,在相同多任务设置下,本方法优于12种对比方法,尤其在数据受限和噪声环境下表现突出。大量实验证明,所学表示更具充分性、数据高效性和鲁棒性。
原文摘要 · Abstract (English)
This paper proposes a new principled multi-task representation learning framework (InfoMTL) to extract noise-invariant sufficient representations for all tasks. It ensures sufficiency of shared representations for all tasks and mitigates the negative effect of redundant features, which can enhance language understanding of pre-trained language models (PLMs) under the multi-task paradigm. Firstly, a shared information maximization principle is proposed to learn more sufficient shared representations for all target tasks. It can avoid the insufficiency issue arising from representation compression in the multi-task paradigm. Secondly, a task-specific information minimization principle is designed to mitigate the negative effect of potential redundant features in the input for each task. It can compress task-irrelevant redundant information and preserve necessary information relevant to the target for multi-task prediction. Experiments on six classification benchmarks show that our method outperforms 12 comparative multi-task methods under the same multi-task settings, especially in data-constrained and noisy scenarios. Extensive experiments demonstrate that the learned representations are more sufficient, data-efficient, and robust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。