模型懂用户问题后仍答不对,关键在理解回复环节。
Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA

- 拆解多轮问答为澄清决策与后续回答两部分
- 澄清策略优化快,但答对率仍低15%以上
- 适合研究对话系统、人机交互的开发者
多元对齐要求系统适应不同用户的价值观、沟通风格和上下文假设。我们认为,准确获取用户偏好是实现对齐的基础,尤其在意图不明确或模糊时。本文通过分解多轮问答中的偏好获取问题,提出两个组件: - 澄清策略(clarification policy):决定是否提问澄清或直接作答; - 澄清后回答(post-clarification answering):在获得缺失信息后生成正确最终答案。 基于PACIFIC基准测试发现,监督微调可快速提升澄清策略性能,但即使模型做出正确澄清动作,最终答案准确率仍显著偏低。该差距表明,理解并正确解析用户回复是多轮问答系统的核心瓶颈。
原文摘要 · Abstract (English)
Pluralistic alignment requires systems to adapt to diverse user values, communication styles, and contextual assumptions. We believe that a foundational prerequisite for such alignment enabling accurate preference elicitation from people when their intent is under-specified or ambiguous. We study the problem of preference elicitation in multi-turn question answering by decomposing the problem into two components: a \textbf{clarification policy}, which decides whether to ask a clarifying question or answer directly, and \textbf{post-clarification answering}, which produces the correct final answer once the missing information is provided. We show, using the PACIFIC benchmark, that supervised fine-tuning rapidly improves the clarification policy, however, final answer accuracy remains substantially lower even when the model takes the correct action. This gap indicates that understanding and correctly interpreting the user's response is the critical gap in multi-turn question-answering systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。