提出新型反馈学习框架,揭示其与传统在线学习的本质差异。
Discriminative Feature Feedback with General Teacher Classes
- 引入判别特征反馈机制,构建可分析的通用学习框架。
- 在可实现和不可实现场景下均给出最优错误率上界。
- 发现传统维度指标无法刻画该框架的泛化能力,适合理论研究者。
我们研究了交互式学习协议判别特征反馈(DFF)的理论性质(Dasgupta et al., 2018)。DFF通过判别性特征解释提供反馈。本文首次在与监督学习和在线学习可比的通用框架下系统分析DFF。研究了可实现与不可实现设定下的最优错误率上界,获得新的结构性结果,并揭示了在线学习与具有更丰富反馈(如DFF)之间的本质差异。在可实现设定中,我们使用一种新提出的维度概念刻画错误率上界;在不可实现设定中,给出了错误率上界并证明其在一般情况下无法改进。结果表明,不同于在线学习,在DFF中,可实现维度不足以刻画最优不可实现错误率或无悔算法的存在性。
原文摘要 · Abstract (English)
We study the theoretical properties of the interactive learning protocol Discriminative Feature Feedback (DFF) (Dasgupta et al., 2018). The DFF learning protocol uses feedback in the form of discriminative feature explanations. We provide the first systematic study of DFF in a general framework that is comparable to that of classical protocols such as supervised learning and online learning. We study the optimal mistake bound of DFF in the realizable and the non-realizable settings, and obtain novel structural results, as well as insights into the differences between Online Learning and settings with richer feedback such as DFF. We characterize the mistake bound in the realizable setting using a new notion of dimension. In the non-realizable setting, we provide a mistake upper bound and show that it cannot be improved in general. Our results show that unlike Online Learning, in DFF the realizable dimension is insufficient to characterize the optimal non-realizable mistake bound or the existence of no-regret algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。