用轨迹信息动态调整拒绝回答的奖励,提升大模型拒答能力。
TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

- 基于策略优化中的多轨迹信息,动态重加权拒答奖励
- 在31个数据集上17次超越基线,5个类别达最优拒答F1
- 适合关注幻觉减少与拒答能力提升的研究者
本文研究大语言模型的拒答学习,提出一种基于轨迹感知的优势重加权方法(TIAR),将传统的三元奖励升级为动态重加权机制。该方法在组相对策略优化(GRPO)训练中,利用多个生成轨迹作为策略对查询的置信度信号,动态计算拒答优势,从而增强模型对知识边界的识别能力。实验采用AbstentionBench基准测试,涵盖六个评估类别、31个数据集。结果表明,TIAR在五个类别上达到当前最佳拒答F1分数,在17个数据集上优于静态三元基线,同时完全保持基线准确率,验证了其在减少幻觉方面的有效性。
原文摘要 · Abstract (English)
This paper investigates large language model (LLM) abstention learning, specifically using ternary reward, which incentivize truthfulness in large language models. This paper extends that idea by moving from a ternary reward to a Trajectory-Informed advantage reweighting, dynamically re-weights the abstention reward during Group Relative Policy Optimization (GRPO) training. The objective of this work focuses on abstention learning instead of improving truthfulness, serving as an exploration into hallucination reduction. The novelty of this paper lies in methodological innovation, advantage re-weighting, and benchmark selection. Leveraging GRPO's multiple trajectories as a natural abstention signal, this method uses a reward signal to explore knowledge boundaries and encourage consistency. By demonstrating that trajectories can be used as a confidence indicator of the policy relative to the query, they are then used to dynamically calculate the abstention advantage. AbstentionBench is used as the evaluation benchmark, as this work aims to contribute to the field of abstention learning. All datasets on the benchmark were tested against this method and various baselines. Empirical results demonstrate that TIAR achieves state-of-the-art abstention F1 scores across five of six evaluation categories, outperforming the static ternary baseline on 17 of 31 benchmark datasets while fully preserving baseline accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。