arXiv:2606.20544cs.AIcs.LG2026-06

提升专家混合模型在分布偏移下的预测可信度。

Toward Calibrated Mixture-of-Experts Under Distribution Shift

论文配图:Toward Calibrated Mixture-of-Experts Under Distribution Shift
图 1 · 摘自论文原文
  • 通过对抗重加权强化路由聚合的校准性。
  • 软路由模型需额外校准机制,硬路由则无需。
  • 适用于多任务、多模型和各类分布偏移场景。

校准使模型预测不确定性与实际发生频率一致,对信任概率输出至关重要。近期研究表明,在个体专家层面强制校准可提升集成模型的准确率与校准性,尤其在专家混合(MoE)模型中表现显著;但校准为何有效仍不清晰。本文研究MoE模型在分布偏移下的行为,聚焦路由机制与专家级校准的交互关系。结果表明:在硬路由模型中,专家校准足以保证整体模型校准;但在软路由模型中则不足。为此,我们提出一种对抗重加权方法,惩罚分布偏移下路由聚合的校准误差,实验证明其在各类模型、任务和分布偏移下均提升了准确率-校准权衡,包括困难数据子集的表现。

原文摘要 · Abstract (English)

Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can improve ensemble accuracy and calibration, with mixture-of-experts (MoE) models showing strong empirical improvements in particular; however, the conditions under which calibration helps MoE are not well understood. In this work, we study how MoE models behave under distribution shift, focusing on how routing mechanisms interact with expert-level calibration. We show that expert calibration is sufficient to ensure calibration of the overall model under a broad class of distribution shifts in hard-routed models, but is insufficient for calibrating soft-routed models. To address this, we propose an adversarial reweighting that penalizes calibration errors of the routed aggregate under distribution shift, and we demonstrate that it improves the accuracy-calibration tradeoff both on average and on difficult subsets of the data, across model classes, prediction tasks, and distribution shifts.

MoE校准分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。