arXiv:2505.16548cs.LGstat.ML2025-05NeurIPS被引 3

通过时间一致性提升序列分类的增量预测效率。

Incremental Sequence Classification with Temporal Consistency

  • 基于强化学习的时间差分思想,设计新损失函数
  • 仅需少量词元即可更好区分生成结果优劣
  • 适用于文本分类与大模型生成验证任务

我们研究增量序列分类问题,即随着序列元素逐步揭示,预测结果随之更新。借鉴强化学习中的时间差分学习思想,我们提出一种时间一致性条件,要求连续预测满足该条件。据此设计了一种新型损失函数,用于训练增量序列分类器。通过具体实例表明,优化该损失可显著提升数据效率。将方法应用于文本分类任务,在多个基准数据集上优于现有方法。进一步在小学数学问题生成正确性验证任务中测试,结果显示使用该方法训练的模型仅观察少数词元后,便能更准确地区分有潜力与无潜力的生成结果。

原文摘要 · Abstract (English)

We address the problem of incremental sequence classification, where predictions are updated as new elements in the sequence are revealed. Drawing on temporal-difference learning from reinforcement learning, we identify a temporal-consistency condition that successive predictions should satisfy. We leverage this condition to develop a novel loss function for training incremental sequence classifiers. Through a concrete example, we demonstrate that optimizing this loss can offer substantial gains in data efficiency. We apply our method to text classification tasks and show that it improves predictive accuracy over competing approaches on several benchmark datasets. We further evaluate our approach on the task of verifying large language model generations for correctness in grade-school math problems. Our results show that models trained with our method are better able to distinguish promising generations from unpromising ones after observing only a few tokens.

序列分类增量学习大模型验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。