让模型在选专家时决定是否加信息,更灵活省钱。
Learning-to-Defer with Expert-Conditional Advice
- 用联合策略空间学习选专家和给信息的组合
- 实验证明比传统方法更省成本,还能自适应调整
- 适合需要动态决策的多模态或语言任务
Learning-to-Defer 通过选择最小化预期成本的专家来处理输入,但假设每个专家在决策时可用的信息是固定的。许多现代系统违背了这一假设:选定专家后,还可决定为其提供额外信息,如检索文档、工具输出或升级上下文。本文研究此问题,称之为带建议的 Learning-to-Defer。我们证明,一类常见的分离式代理(路由与建议使用独立头部)在最简非平凡设置下即存在不一致。随后提出一种扩展型代理,作用于专家-建议复合动作空间,并建立 $/mathcal{H}$-一致性保证与过失风险转移界,确保极限情况下恢复贝叶斯最优策略。在表格、语言及多模态任务上的实验表明,该方法优于标准 Learning-to-Defer,且能根据成本环境自适应调整建议获取行为;合成基准验证了分离式代理的失效模式。
原文摘要 · Abstract (English)
Learning-to-Defer routes each input to the expert that minimizes expected cost, but it assumes that the information available to every expert is fixed at decision time. Many modern systems violate this assumption: after selecting an expert, one may also choose what additional information that expert should receive, such as retrieved documents, tool outputs, or escalation context. We study this problem and call it Learning-to-Defer with advice. We show that a broad family of natural separated surrogates, which learn routing and advice with distinct heads, is inconsistent even in the smallest non-trivial setting. We then introduce an augmented surrogate that operates on the composite expert--advice action space and prove an $\mathcal{H}$-consistency guarantee together with an excess-risk transfer bound, yielding recovery of the Bayes-optimal policy in the limit. Experiments on tabular, language, and multi-modal tasks show that the resulting method improves over standard Learning-to-Defer while adapting its advice-acquisition behavior to the cost regime; a synthetic benchmark confirms the failure mode predicted for separated surrogates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。