arXiv:2501.12115cs.LGcs.CV2025-01

用元学习自动学出多任务网络的最佳稀疏结构。

Meta-Sparsity: Learning Optimal Sparse Structures in Multi-task Networks through Meta-learning

  • 通过元学习动态优化各任务的稀疏模式。
  • 在NYU-v2和CelebAMask-HQ上跨任务表现优异。
  • 适合追求高效可扩展模型的研究者。

本文提出Meta-Sparsity框架,通过元学习自动学习控制稀疏度的参数,使深度神经网络在多任务学习中自动生成最优稀疏共享结构。与依赖人工调参的传统稀疏方法不同,该方法受MAML启发,在元训练阶段引入基于惩罚的通道级结构化稀疏,实现对共享稀疏参数的学习。该方法有效去除冗余参数,提升模型在已见及未见任务上的泛化能力。在NYU-v2和CelebAMask-HQ两个数据集上进行的广泛实验表明,该方法在像素级到图像级预测等多种任务中表现稳健,验证了其作为高效、可适应稀疏神经网络通用工具的潜力。本工作为稀疏神经网络研究提供了新方向。

原文摘要 · Abstract (English)

This paper presents meta-sparsity, a framework for learning model sparsity, basically learning the parameter that controls the degree of sparsity, that allows deep neural networks (DNNs) to inherently generate optimal sparse shared structures in multi-task learning (MTL) setting. This proposed approach enables the dynamic learning of sparsity patterns across a variety of tasks, unlike traditional sparsity methods that rely heavily on manual hyperparameter tuning. Inspired by Model Agnostic Meta-Learning (MAML), the emphasis is on learning shared and optimally sparse parameters in multi-task scenarios by implementing a penalty-based, channel-wise structured sparsity during the meta-training phase. This method improves the model's efficacy by removing unnecessary parameters and enhances its ability to handle both seen and previously unseen tasks. The effectiveness of meta-sparsity is rigorously evaluated by extensive experiments on two datasets, NYU-v2 and CelebAMask-HQ, covering a broad spectrum of tasks ranging from pixel-level to image-level predictions. The results show that the proposed approach performs well across many tasks, indicating its potential as a versatile tool for creating efficient and adaptable sparse neural networks. This work, therefore, presents an approach towards learning sparsity, contributing to the efforts in the field of sparse neural networks and suggesting new directions for research towards parsimonious models.

多任务学习稀疏网络元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。