arXiv:2602.09402cs.LG2026-02被引 2

研究多个正确答案下的学习,揭示不同反馈机制的错误率边界。

Learning with Multiple Correct Answers -- Regret Bounds under Different Feedback Models

  • 针对多正确答案场景设计在线学习算法
  • 三种反馈下后悔率分别为常数、线性或次线性
  • 适用于自然语言生成等开放任务

我们研究具有多个正确答案的学习问题,其中每个实例允许多个有效标签。主要关注在线设置,每轮学习者需对查询样本输出一个有效标签。该设定源于语言生成任务,即提示可能有多种合理补全,但并非所有补全都可接受。我们在三种反馈模型下研究此问题。对每种模型,利用适当的组合维度刻画可实现设置下的最优误判边界,并在广义设置下表明后悔率可为常数、线性或次线性。结果还推导出批量学习下的样本复杂度界,其依赖于相应的组合维度。

原文摘要 · Abstract (English)

We study the problem of learning with multiple correct answers, where each instance admits a set of valid labels. We primarily focus on the online setup, where in each round the learner must output a valid label for the queried example. This setting is motivated by language generation, in which a prompt may admit many acceptable completions, but not every completion is acceptable. We study this problem under three feedback models. For each model, we characterize the optimal mistake bound in the realizable setting using an appropriate combinatorial dimension. We then show that the rate of regret can be constant, linear, or sublinear across the three models in the agnostic setting. Our results also imply sample complexity bounds for the batch setup that depend on the respective combinatorial dimensions.

在线学习多正确答案后悔率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。