对比人类与大模型在任务对话中对指代模糊的澄清行为差异
Referential ambiguity and clarification requests: comparing human and LLM behaviour
- 构建统一标注数据集,整合指代模糊与澄清请求信息
- 人类极少因指代模糊提问,却常为任务不确定性提问;大模型反之
- 推理能力提升可增强大模型澄清问题的频率与相关性
本研究考察大模型在异步指令发送-执行对话中提出澄清问题的能力。我们构建了一个新语料库,将Minecraft Dialogue Corpus中的指代与模糊标注、以及基于SDRT的澄清标注整合为统一格式,支持澄清行为与模糊性的关联研究。通过该语料库,我们对比了人类与大模型在模糊情境下的行为。结果发现,人类产生澄清问题与模糊性之间关联较弱,且人类极少针对指代模糊提问,但常因任务不确定性提问;而大模型更倾向于对指代模糊提问,对任务不确定性则较少。我们质疑大模型的澄清能力是否源于其近期具备的模拟推理能力,并测试不同推理方法,发现推理确实能提高大模型提问的频率和相关性。
原文摘要 · Abstract (English)
In this work we examine LLMs' ability to ask clarification questions in task-oriented dialogues that follow the asynchronous instruction-giver/instruction-follower format. We present a new corpus that combines two existing annotations of the Minecraft Dialogue Corpus -- one for reference and ambiguity in reference, and one for SDRT including clarifications -- into a single common format providing the necessary information to experiment with clarifications and their relation to ambiguity. With this corpus we compare LLM actions with original human-generated clarification questions, examining how both humans and LLMs act in the case of ambiguity. We find that there is only a weak link between ambiguity and humans producing clarification questions in these dialogues, and low correlation between humans and LLMs. Humans hardly ever produce clarification questions for referential ambiguity, but often do so for task-based uncertainty. Conversely, LLMs produce more clarification questions for referential ambiguity, but less so for task uncertainty. We question if LLMs' ability to ask clarification questions is predicated on their recent ability to simulate reasoning, and test this with different reasoning approaches, finding that reasoning does appear to increase question frequency and relevancy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。