arXiv:2602.17144cs.LGstat.ML2026-02被引 2

多专家模型反而更易欠拟合,新方法通过选可信专家解决此问题。

When More Experts Hurt: Underfitting in Multi-Expert Learning to Defer

  • 提出PiCCE方法,根据证据自适应选择可靠专家
  • 实验证明可显著提升多专家学习的预测性能
  • 适合需要信任专家系统的高可靠性场景

学习去延迟(L2D)使分类器能放弃预测并转交专家处理,最近已扩展至多专家设置。本文揭示,多专家L2D比单专家情形更具挑战性:分类器的欠拟合成为固有现象,严重损害预测性能,而单专家情况下仅在特定条件下出现。理论分析表明,这源于内在的专家可辨识性问题——从多样化专家池中学习信任谁,该问题在单专家情况下不存在,导致现有欠拟合缓解方法失效。为此,我们提出基于代理的PiCCE(选自信且正确的专家)方法,通过经验证据自适应识别可靠专家,将多专家L2D转化为类单专家学习问题,从而解决多专家欠拟合。进一步证明了其统计一致性,以及恢复类别概率和专家准确率的能力。在多种设置(包括真实世界专家场景)上的大量实验验证了理论结果,并展示了性能提升。

原文摘要 · Abstract (English)

Learning to Defer (L2D) enables a classifier to abstain from predictions and defer to an expert, and has recently been extended to multi-expert settings. In this work, we show that multi-expert L2D is fundamentally more challenging than the single-expert case. With multiple experts, the classifier's underfitting becomes inherent, which seriously degrades prediction performance, whereas in the single-expert setting it arises only under specific conditions. We theoretically reveal that this stems from an intrinsic expert identifiability issue: learning which expert to trust from a diverse pool, a problem absent in the single-expert case and renders existing underfitting remedies failed. To tackle this issue, we propose PiCCE (Pick the Confident and Correct Expert), a surrogate-based method that adaptively identifies a reliable expert based on empirical evidence. PiCCE effectively reduces multi-expert L2D to a single-expert-like learning problem, thereby resolving multi expert underfitting. We further prove its statistical consistency and ability to recover class probabilities and expert accuracies. Extensive experiments across diverse settings, including real-world expert scenarios, validate our theoretical results and demonstrate improved performance.

多专家学习去延迟欠拟合专家选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。