arXiv:2602.00939math.STcs.LG2026-02

首次分析分类任务中专家异质性对参数估计效率的影响。

Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity

  • 构建异质与同质结构的多元逻辑回归专家混合模型,分析其收敛性。
  • 在参数随样本量变化的挑战设定下,获得统一收敛率并证明其极小极大最优。
  • 揭示专家异质性可提升估计速度,更高效利用数据,适合迁移学习场景。

污染型专家混合(MoE)模型源于迁移学习,其中预训练的冻结专家与可训练适配器结合以学习新任务。尽管已有研究分析该模型的参数估计收敛性,但文献仍存在两大空白:其一,污染型MoE仅在回归设置中被研究,其在分类任务中的理论基础缺失;其二,现有分类模型虽有参数估计的逐点收敛率,却缺乏极小极大最优性的保证。本文首次对具有同质与异质结构的污染型多元逻辑回归专家混合模型进行收敛性分析。在每种结构下,刻画了在真实参数随样本量变化的困难设定中参数估计的统一收敛率,并建立了相应的极小极大下界以证明其最优性。理论表明,专家异质性可带来更快的参数估计速率,因而比同质结构更具样本效率。

原文摘要 · Abstract (English)

Contaminated mixture of experts (MoE) is motivated by transfer learning methods where a pre-trained model, acting as a frozen expert, is integrated with an adapter model, functioning as a trainable expert, in order to learn a new task. Despite recent efforts to analyze the convergence behavior of parameter estimation in this model, there are still two unresolved problems in the literature. First, the contaminated MoE model has been studied solely in regression settings, while its theoretical foundation in classification settings remains absent. Second, previous works on MoE models for classification capture pointwise convergence rates for parameter estimation without any guaranty of minimax optimality. In this work, we close these gaps by performing, for the first time, the convergence analysis of a contaminated mixture of multinomial logistic experts with homogeneous and heterogeneous structures, respectively. In each regime, we characterize uniform convergence rates for estimating parameters under challenging settings where ground-truth parameters vary with the sample size. Furthermore, we also establish corresponding minimax lower bounds to ensure that these rates are minimax optimal. Notably, our theories offer an important insight into the design of contaminated MoE, that is, expert heterogeneity yields faster parameter estimation rates and, therefore, is more sample-efficient than expert homogeneity.

专家混合分类模型迁移学习极小极大最优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。