在专家资源有限时,动态决定何时由模型自动决策,何时交给人类。
Online Decision Deferral under Budget Constraints
- 用上下文老虎机框架建模在线决策问题,支持预算约束
- 算法在真实数据集上表现优异,能适应任务分布变化
- 适合需要持续优化人机协作的高成本决策场景
机器学习模型日益用于支持或替代决策。在专家资源有限的应用中,减少其负担并自动化决策至关重要,前提是模型性能不低于人工水平。然而,模型常为预训练固定,而任务按序到达且分布可能漂移,导致决策者表现变化,因此需自适应的延迟策略。本文提出一种包含预算约束和多种部分反馈机制的上下文老虎机模型。除理论保证外,还设计高效扩展,在真实数据集上取得显著性能。
原文摘要 · Abstract (English)
Machine Learning (ML) models are increasingly used to support or substitute decision making. In applications where skilled experts are a limited resource, it is crucial to reduce their burden and automate decisions when the performance of an ML model is at least of equal quality. However, models are often pre-trained and fixed, while tasks arrive sequentially and their distribution may shift. In that case, the respective performance of the decision makers may change, and the deferral algorithm must remain adaptive. We propose a contextual bandit model of this online decision making problem. Our framework includes budget constraints and different types of partial feedback models. Beyond the theoretical guarantees of our algorithm, we propose efficient extensions that achieve remarkable performance on real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。