让大模型学会准确拒绝回答并说明缺了什么信息
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL

- 用可验证的强化学习奖励,同时优化拒绝与澄清
- 3B模型在无法回答的问题上表现优于基线,接近更大模型
- 适合需要可靠拒答和清晰解释的场景,如医疗咨询
强化微调提升了大语言模型的推理能力,但也可能诱使模型在缺乏信息时猜测或幻觉。现有拒答方法要么训练模型产生通用拒绝,要么鼓励追问但不验证追问是否指出了关键缺失。我们研究语义清晰但无法从给定信息中可靠解决的问题,认为可靠模型不仅应拒答,还应说明缺失内容。提出一种澄清感知的可验证强化学习奖励(RLVR),在奖励可回答问题正确回答的同时,联合优化不可回答问题上的显式拒答与语义对齐的拒答后澄清。基于该奖励训练出的3B参数量模型Abstain-R1,在不可回答问题上的拒答与澄清能力显著提升,同时保持可回答问题的强性能。在Abstain-Test、Abstain-QA和SelfAware数据集上的实验表明,Abstain-R1明显优于其基线模型,并在不可回答查询行为上达到与DeepSeek-R1等更大系统相当的水平,表明校准的拒答与澄清可通过可验证奖励学习,而非仅依赖模型规模。
原文摘要 · Abstract (English)
Reinforcement fine-tuning improves the reasoning ability of large language models, but it can also encourage them to answer unanswerable queries by guessing or hallucinating missing information. Existing abstention methods either train models to produce generic refusals or encourage follow-up clarifications without verifying whether those clarifications identify the key missing information. We study queries that are clear in meaning but cannot be reliably resolved from the given information, and argue that a reliable model should not only abstain, but also explain what is missing. We propose a clarification-aware RLVR reward that, while rewarding correct answers on answerable queries, jointly optimizes explicit abstention and semantically aligned post-refusal clarification on unanswerable queries. Using this reward, we train Abstain-R1, a 3B model that improves abstention and clarification on unanswerable queries while preserving strong performance on answerable ones. Experiments on Abstain-Test, Abstain-QA, and SelfAware show that Abstain-R1 substantially improves over its base model and achieves unanswerable-query behavior competitive with larger systems including DeepSeek-R1, suggesting that calibrated abstention and clarification can be learned through verifiable rewards rather than emerging from scale alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。