arXiv:2602.11199cs.CLcs.LG2026-02ACL被引 5

教大模型在该问时问,问对问题,避免胡说八道。

When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification

  • 用交互式评测框架让模型学会判断何时该提问
  • 提问准确率提升,且对话更高效,跨领域泛化强
  • 适合想提升模型可靠性的研究者和开发者

大型语言模型常在关键信息缺失或存在误导性内容时仍强行回答,导致幻觉或强化错误认知。本文研究如何评估并提升模型在何时、何事上应提出澄清的问题,同时不牺牲任务表现。提出 AskBench 交互式评测基准,将标准问答对转化为带明确检查点的多轮互动;统一评判循环评估最终答案,并按需模拟用户反馈。覆盖两种场景:AskMind(意图不全的提问需澄清)与 AskOverconfidence(包含错误前提需识别纠正)。进一步提出基于评分表引导的验证器奖励强化学习(RLVR),利用结构化评分表促进精准提问。实验表明,在准确性、评分表符合度和交互效率上均有稳定提升,且在未见领域中具有良好泛化能力。

原文摘要 · Abstract (English)

Large language models (LLMs) often respond even when prompts omit critical details or include misleading information, leading to hallucinations or reinforced misconceptions. We study how to evaluate and improve LLMs' ability to decide when and what to ask for clarification without sacrificing task performance. We introduce AskBench, an interactive benchmark that converts standard QA pairs into multi-turn interactions with explicit checkpoints. A unified judge loop evaluates final answers and simulates user responses as needed. AskBench covers two settings: AskMind, with intent-deficient queries requiring clarification, and AskOverconfidence, with queries containing false premises that must be identified and corrected. We further propose rubric-guided reinforcement learning with verifier-based rewards (RLVR), which uses structured rubrics to encourage targeted clarification. Experiments show consistent improvements in accuracy, rubric adherence, and interaction efficiency, with strong generalization to unseen domains.

大模型澄清提问强化学习评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。