arXiv:2506.20650cs.LGstat.ML2025-06ICML被引 23

提出新型代理损失函数,实现多专家延迟学习的强理论保证

Mastering Multiple-Expert Routing: Realizable $H$-Consistency and Strong Guarantees for Learning to Defer

  • 设计可实现的H-一致性代理损失,优化多专家延迟决策
  • 在单阶段与双阶段场景下均获得理论保证,涵盖贝叶斯一致性
  • 适用于语言生成、医疗诊断等需权衡准确率与计算成本的场景

多专家延迟学习的核心是优化输入实例分配策略,平衡专家的准确性与计算开销。该问题在自然语言生成、图像处理和医学诊断中至关重要。尽管已有研究提出代理损失函数来优化延迟策略,但其一致性性质仍存挑战。本文提出新的代理损失函数及高效算法,在单阶段(联合学习预测器与延迟函数)和双阶段(固定专家仅学延迟函数)场景下,解决了可实现H-一致性、H-一致性界和贝叶斯一致性等开放问题。针对单阶段,提出一类可实现H-一致的代理损失,并证明其中一成员具备H-一致性。针对双阶段,推导出在两专家场景下实现可实现H-一致性、一致性界与贝叶斯一致性的新损失,并在合理假设下扩展至多专家场景。此外,在低噪声假设下进一步提升理论保证。实验验证了所提方法优于现有基线。

原文摘要 · Abstract (English)

The problem of learning to defer with multiple experts consists of optimally assigning input instances to experts, balancing the trade-off between their accuracy and computational cost. This is a critical challenge in natural language generation, but also in other fields such as image processing, and medical diagnostics. Recent studies have proposed surrogate loss functions to optimize deferral, but challenges remain in ensuring their consistency properties. This paper introduces novel surrogate loss functions and efficient algorithms with strong theoretical learning guarantees. We address open questions regarding realizable $H$-consistency, $H$-consistency bounds, and Bayes-consistency for both single-stage (jointly learning predictor and deferral function) and two-stage (learning only the deferral function with a fixed expert) learning scenarios. For single-stage deferral, we introduce a family of new realizable $H$-consistent surrogate losses and further prove $H$-consistency for a selected member. For two-stage deferral, we derive new surrogate losses that achieve realizable $H$-consistency, $H$-consistency bounds, and Bayes-consistency for the two-expert scenario and, under natural assumptions, multiple-expert scenario. Additionally, we provide enhanced theoretical guarantees under low-noise assumptions for both scenarios. Finally, we report the results of experiments using our proposed surrogate losses, comparing their performance against existing baselines.

延迟学习多专家理论保证代理损失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。