arXiv:2607.21143cs.CLcs.AI2026-07

评估大模型澄清策略的多轮互动基准,看它问得对不对、何时停。

One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

  • 用后悔值衡量模型澄清策略的优劣,而非单轮问题质量。
  • 相同准确率下,不同模型在效率和停止时机上差异显著。
  • 适合研究对话系统优化与用户行为鲁棒性的研究者。

模糊用户请求使澄清成为对话式大模型助手的序列决策问题:需决定是否提问、问什么、何时停止、何时回答。我们提出RegretBench,一个将澄清视为策略行为而非孤立问题质量的多轮评测基准。该基准采用隐含意图的模糊性建模,支持基于语义状态跟踪的自由交互,并引入基于后悔值的目标,衡量模型相对于参考澄清策略所损失的价值。在开放域问答与产品推荐场景的实验表明,最终成功率不足以评估模型表现,因准确率相近的模型在效率、对用户行为的鲁棒性及停止决策方面存在显著差异。通过联合测量意图解析、交互成本、无效澄清和后悔值,RegretBench揭示模型是否以有效且高效的方式进行澄清。结果表明,有效的澄清不仅需要合理的问题,更需在正确时机提问并在用户意图明确后及时停止。

原文摘要 · Abstract (English)

Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to answer. We introduce RegretBench, a multi-turn benchmark that evaluates clarification as policy behavior rather than isolated question quality. RegretBench provides a hidden-intent formulation of ambiguity, supports free-form interaction grounded in semantic-state tracking, and introduces a regret-based objective that measures how much value a model loses relative to a reference clarification policy. Experiments on open-domain QA and product recommendation scenarios show that final success alone is insufficient, as models with similar accuracy can differ substantially in efficiency, robustness to user behaviors, and stopping decisions. By jointly measuring intent resolution, interaction cost, ineffective clarification, and regret, RegretBench reveals whether models clarify usefully and efficiently. Our results show that effective clarification requires more than plausible questions: models must ask the right question at the right time and stop once the user's intended meaning is clear.

对话系统澄清策略多轮交互评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。