提出新方法提升视觉多任务学习的稳定性与性能
SGW-based Multi-Task Learning in Vision Tasks
- 通过信息瓶颈模块抑制任务间干扰
- 在多个数据集上优于现有方法,显著提升多任务表现
- 适合研究多任务学习与模型优化的开发者
多任务学习(MTL)是一种多目标优化任务,神经网络在共享表示空间中尝试实现各个目标。然而,随着数据集规模扩大和任务复杂度增加,知识共享变得愈发困难。本文从噪声角度重新审视以往基于交叉注意力的MTL方法,理论分析发现其机制存在缺陷。为此,我们提出信息瓶颈知识提取模块(KEM),通过约束信息流减少任务间干扰,降低计算复杂度。此外,引入神经坍缩思想,在输入KEM前将特征投影至ETF空间,以稳定知识选择过程。我们在多个数据集上实现了该方法并进行了对比实验,结果表明,该方法在多任务学习中显著优于现有方法。
原文摘要 · Abstract (English)
Multi-task-learning(MTL) is a multi-target optimization task. Neural networks try to realize each target using a shared interpretative space within MTL. However, as the scale of datasets expands and the complexity of tasks increases, knowledge sharing becomes increasingly challenging. In this paper, we first re-examine previous cross-attention MTL methods from the perspective of noise. We theoretically analyze this issue and identify it as a flaw in the cross-attention mechanism. To address this issue, we propose an information bottleneck knowledge extraction module (KEM). This module aims to reduce inter-task interference by constraining the flow of information, thereby reducing computational complexity. Furthermore, we have employed neural collapse to stabilize the knowledge-selection process. That is, before input to KEM, we projected the features into ETF space. This mapping makes our method more robust. We implemented and conducted comparative experiments with this method on multiple datasets. The results demonstrate that our approach significantly outperforms existing methods in multi-task learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。