共享特征能显著减少多任务模型的描述长度,尤其在正交约束下。
The Information-Theoretic Benefit of Shared Representations under Orthogonality Constraints

- 用正交结构构造共享特征与任务专用读出,实现联合逼近
- 理论证明联合编码比独立编码少用比特,提升描述效率
- 适用于需几何约束的多任务神经网络设计
现代深度学习架构日益趋向多任务、多模态,常以预训练基础模型结合任务特定微调模型。实证表明,利用不同任务间的相似性而非单独求解,可显著提升整体性能。尽管多任务学习的泛化性和样本复杂性已有广泛研究,但联合逼近与独立逼近在参数复杂度上的差异仍不明确。本文在一致范数下证明了独立与联合逼近类别的描述长度上下界。通过构造一类正交函数,将共享硬特征(由Rademacher-Haar小波级数实现)与锯齿-沃尔什读出结合,强制输出坐标的正交性。小波的二叉树结构使逼近难度集中在共享特征部分,而读出则作为任务专用头。基于信息论框架,我们得到联合与独立编码最优逼近率之间的显著差距。最终,通过将赫维赛德激活的神经网络简化为三角波逼近,实现了该分离。结果表明,在正交约束下,只要任务共享潜在硬特征,联合逼近所需比特数严格少于独立逼近,为组合型多输出架构的描述长度效率提供了理论支持,并阐明了神经网络在几何约束下如何保持表达能力。
原文摘要 · Abstract (English)
Modern deep learning architectures are increasingly multi-task and multi-modal, using a pretrained foundation model combined with task-specific, fine-tuned models. Empirically, exploiting similarity across different problems, instead of solving them individually, can significantly improve overall performance. While the generalization and sample complexity properties of multitask learning have been widely studied, the parametric complexity of joint approximation in comparison to separate approximation remains less well understood. The question is particularly relevant in modern deep learning, where models are increasingly required to satisfy structural constraints such as equivariance, conservation laws, or orthogonality. We prove lower and upper bounds on the description-length for separate and joint approximation classes, respectively, in uniform norm. We build a class of orthogonal functions by composing a shared hard feature, realized by a Rademacher-Haar wavelet series, with Sawtooth-Walsh readouts to enforce orthogonality of output coordinates. The dyadic tree structure of the Rademacher-Haar wavelet concentrates the approximation hardness in the common feature component, while the readouts act as task-specific heads. Using an information-theoretic framework, we obtain a sharp gap between the optimal approximation rates achievable by joint and separate coding. Finally, we realize this separation in a neural network model using Heaviside activations via reduction to triangle-wave approximation. Our results show that even under an orthogonality constraint joint approximation requires strictly fewer bits in compositional architectures, provided the tasks share a latent hard feature. This provides theoretical insight into the description-length-efficiency of compositional multi-output architectures and clarifies how neural networks can retain expressivity under geometric constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。