解析提示词微调中专家模型的收敛问题,揭示参数消失与交互阻碍机制。
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
- 引入可区分性条件控制预训练模型与提示词间的参数干扰
- 证明不同专家结构下参数估计的收敛速率及最优下界
- 适用于研究大模型微调机制的理论工作者
我们对污染混合专家模型中的参数估计收敛性进行了分析。该模型源于提示词微调问题,其中提示词可视为专家,用于微调大规模预训练模型以学习下游任务。分析中浮现两个根本挑战:(i) 预训练模型与提示词在混合比例中可能收敛至零,导致提示词消失问题;(ii) 预训练模型与提示词参数间通过偏微分方程产生代数交互,减缓提示词学习速度。为此,我们引入可区分性条件以控制前述参数交互。此外,还研究了多种专家结构对参数估计收敛行为的影响。在每种情形下,均给出了参数估计的完整收敛速率及相应的极小极大下界。最后,通过多个数值实验验证了理论结果的正确性。
原文摘要 · Abstract (English)
We conduct the convergence analysis of parameter estimation in the contaminated mixture of experts. This model is motivated from the prompt learning problem where ones utilize prompts, which can be formulated as experts, to fine-tune a large-scale pre-trained model for learning downstream tasks. There are two fundamental challenges emerging from the analysis: (i) the proportion in the mixture of the pre-trained model and the prompt may converge to zero during the training, leading to the prompt vanishing issue; (ii) the algebraic interaction among parameters of the pre-trained model and the prompt can occur via some partial differential equations and decelerate the prompt learning. In response, we introduce a distinguishability condition to control the previous parameter interaction. Additionally, we also investigate various types of expert structure to understand their effects on the convergence behavior of parameter estimation. In each scenario, we provide comprehensive convergence rates of parameter estimation along with the corresponding minimax lower bounds. Finally, we run several numerical experiments to empirically justify our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。