arXiv:2608.24195stat.MLcs.LG2026-08

让专家模型自适应选择不同学习方式,提升可解释性与预测性能。

A Heterogeneous Mixture of Experts Framework for Interpretable Machine Learning

  • 引入决策树、SVM、二次判别分析等异质专家,按数据局部特征自动匹配
  • 在多个真实与合成数据集上,性能媲美同质专家模型和随机森林
  • 理论保证优化过程单调上升,且所有专家输出概率一致,便于解释

Mixture-of-Experts(MoE)模型通过输入依赖的门控机制将复杂预测问题分解为更简单的局部学习任务。现有可解释的MoE方法如混合决策树(MoDT)采用同质决策树专家,限制了全特征空间的单一归纳偏置。本文扩展MoDT框架,引入包含决策树、线性支持向量机和二次判别分析的异质专家族,在统一的概率门控机制下运行。为保证基于似然的一致推断,非概率专家经校准以输出条件类别概率,使参数估计可在广义期望最大化(EM)框架内完成。进一步建立了异质门控更新的理论单调上升保证,为优化过程提供依据。在多样化的合成与真实世界基准数据集上的实验表明,该框架能根据局部数据几何结构自适应地分配专家,实现可解释的专家指派,同时预测性能与同质MoDT及随机森林相当。所提方法在统一框架内融合可解释性、自适应归纳偏置选择与概率一致性。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) models provide a flexible framework for partitioning complex prediction problems into simpler local learning tasks through an input-dependent gating mechanism. Existing interpretable MoE approaches, such as Mixture of Decision Trees (MoDT), achieve transparency by employing homogeneous decision-tree experts, but this restricts the model to a single inductive bias across all regions of the feature space. We extend the MoDT framework by introducing heterogeneous expert families comprising decision trees, linear support vector machines, and quadratic discriminant analysis under a common probabilistic gating mechanism. To ensure coherent likelihood-based inference, non-probabilistic experts are calibrated to produce conditional class probabilities, allowing parameter estimation within the generalized Expectation-Maximization framework of MoDT. We further establish theoretical monotone ascent guarantees for the proposed heterogeneous gating updates, providing a justification for the optimization procedure. Experiments on a diverse collection of synthetic and real-world benchmark datasets demonstrate that the proposed framework adaptively specializes experts according to local data geometry, yielding interpretable expert assignments while achieving predictive performance competitive with homogeneous MoDT and Random Forests. The proposed approach combines interpretability, adaptive inductive bias selection, and probabilistic coherence within a unified mixture-of-experts framework.

可解释模型专家混合异质专家概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。