用分布几何结构融合多领域多任务专家模型,提升通用模型性能。
Hierarchical Wasserstein Merging for Multi-Domain Multi-Task Learning: From Specialists to a Generalist

- 在表示空间构建层级最优传输中心,捕捉域与任务间分布关系。
- 在四个领域四类NLP任务上,泛化能力优于现有方法。
- 适合需要跨领域迁移的通用模型研究者使用。
多领域多任务学习(MD-MTL)旨在构建一个能在异构领域和任务上表现良好的通用模型。然而,联合训练常因分布偏移导致干扰。现有模型融合方法多基于参数层面,忽视了不同领域和任务间隐表示分布的几何结构。为此,我们提出分层最优传输融合(HWM),一种表示级框架,将每个领域-任务专家建模为共享支撑上的隐藏表示分布。HWM构建任务级和全局最优传输巴氏中心,以捕捉任务内域差异和跨任务结构,支持无需训练的专家加权聚合,或通过混合最优传输对齐损失实现基于训练的通用模型学习。在每项任务下涵盖四个领域的四项NLP任务上的实验表明,HWM在MD-MTL设置中展现出更优的有效性和泛化能力。
原文摘要 · Abstract (English)
Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across heterogeneous domains and tasks. However, joint training often suffers from interference under distribution shifts. Existing model merging methods mostly operate on model parameters while overlooking the geometric structure of latent representation distributions across domains and tasks. To address these limitations, we propose Hierarchical Wasserstein Merging (HWM), a representation-level framework that models each domain-task specialist as a distribution of hidden representations on a shared support. HWM constructs task-level and global Wasserstein barycenters to capture within-task domain variation and cross-task structure, enabling either training-free specialist aggregation by Wasserstein-derived weights or training-based generalist learning through a hybrid Wasserstein alignment loss. Experiments on four NLP tasks across four domains per task show that HWM achieves superior effectiveness and generalization capability in MD-MTL settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。