通过融合专家分布提升不确定性量化,避免传统模型的多峰伪影。
CoCoAFusE: Beyond Mixtures of Experts via Model Fusion
- 在专家混合基础上引入分布融合机制,增强建模灵活性。
- 相比经典MoE,可信区间更紧致,减少平滑过渡中的伪多峰现象。
- 适合关注不确定性量化与可解释性的复杂回归任务研究者。
许多学习问题涉及多种模式及随协变量变化的不确定性。深度学习虽能捕捉非线性输入输出关系,但模型可解释性与不确定性量化(UQ)进展滞后。本文提出竞争/协作专家融合(CoCoAFusE),一种基于贝叶斯协变量依赖建模的新方法。该方法延续混合专家(MoE)思想,通过多个简单子模型(专家)的预测融合实现高表达力并保持局部可解释性。其创新在于不仅对专家输出进行加权混合,还融合其概率分布,从而更好刻画生成机制间的中间行为,获得更紧致的响应变量可信区间。仅使用传统混合易在平滑过渡处产生多峰伪影,而CoCoAFusE即使在相同结构与先验下亦可避免此问题,显著提升表达力与适应性。在一系列数值示例与真实数据集上验证了其在复杂回归中处理不确定性的有效性。
原文摘要 · Abstract (English)
Many learning problems involve multiple patterns and varying degrees of uncertainty dependent on the covariates. Advances in Deep Learning (DL) have addressed these issues by learning highly nonlinear input-output dependencies. However, model interpretability and Uncertainty Quantification (UQ) have often straggled behind. In this context, we introduce the Competitive/Collaborative Fusion of Experts (CoCoAFusE), a novel, Bayesian Covariates-Dependent Modeling technique. CoCoAFusE builds on the very philosophy behind Mixtures of Experts (MoEs), blending predictions from several simple sub-models (or "experts") to achieve high levels of expressiveness while retaining a substantial degree of local interpretability. Our formulation extends that of a classical Mixture of Experts by contemplating the fusion of the experts' distributions in addition to their more usual mixing (i.e., superimposition). Through this additional feature, CoCoAFusE better accommodates different scenarios for the intermediate behavior between generating mechanisms, resulting in tighter credible bounds on the response variable. Indeed, only resorting to mixing, as in classical MoEs, may lead to multimodality artifacts, especially over smooth transitions. Instead, CoCoAFusE can avoid these artifacts even under the same structure and priors for the experts, leading to greater expressiveness and flexibility in modeling. This new approach is showcased extensively on a suite of motivating numerical examples and a collection of real-data ones, demonstrating its efficacy in tackling complex regression problems where uncertainty is a key quantity of interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。