arXiv:2409.19075cs.CLcs.AI2024-09

动态加权源任务提升低资源常识推理效果

Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning

  • 用强化学习自动评估各源任务对目标任务的贡献度
  • 在极端低资源下性能超越强基线,提升显著
  • 适合常识推理、小样本学习等场景的研究者

元学习广泛用于利用高资源源任务提升低资源目标任务的性能。然而,现有方法通常平等对待不同源任务,忽视了它们与目标任务的相关性。为此,我们提出一种基于强化学习的多源元迁移学习框架(Meta-RTL),用于低资源常识推理。该框架通过强化学习动态估计源任务权重,衡量其在元迁移学习中对目标任务的贡献。将元模型的通用损失与源任务特定时序元模型在采样目标数据上的任务特定损失差异,作为奖励输入强化学习模块的策略网络。策略网络基于LSTM构建,能捕捉跨元学习迭代的长期依赖关系。我们在三个常识推理基准数据集上,使用BERT和ALBERT作为元模型骨干进行评估。实验结果表明,Meta-RTL显著优于强基线及以往任务选择策略,在极低资源设置下提升更大。

原文摘要 · Abstract (English)

Meta learning has been widely used to exploit rich-resource source tasks to improve the performance of low-resource target tasks. Unfortunately, most existing meta learning approaches treat different source tasks equally, ignoring the relatedness of source tasks to the target task in knowledge transfer. To mitigate this issue, we propose a reinforcement-based multi-source meta-transfer learning framework (Meta-RTL) for low-resource commonsense reasoning. In this framework, we present a reinforcement-based approach to dynamically estimating source task weights that measure the contribution of the corresponding tasks to the target task in the meta-transfer learning. The differences between the general loss of the meta model and task-specific losses of source-specific temporal meta models on sampled target data are fed into the policy network of the reinforcement learning module as rewards. The policy network is built upon LSTMs that capture long-term dependencies on source task weight estimation across meta learning iterations. We evaluate the proposed Meta-RTL using both BERT and ALBERT as the backbone of the meta model on three commonsense reasoning benchmark datasets. Experimental results demonstrate that Meta-RTL substantially outperforms strong baselines and previous task selection strategies and achieves larger improvements on extremely low-resource settings.

元学习常识推理低资源强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。