测试大模型在协作任务中能否根据语境调整指令理解,发现模型知错不改。
When Contextual Inference Fails: Cancelability in Interactive Instruction Following
- 设计互动任务让模型判断是否应追问澄清
- 模型虽能识别说话人不可靠但仍盲目猜测
- 适合研究人机协作中推理与行动脱节问题
我们研究了在协作积木搭建任务中,模型如何区分字面理解与语境推断。将双说话人心理语言学范式改编为名为Build What I Mean(BWIM)的交互基准,对比讲究语用合作的说话人与仅字面可靠的说话人。在BWIM中,模型面对模糊指令,需权衡进行语境推断或以小代价请求澄清。评估多个领先LLM后发现,模型在判断与行动间存在明显分离:尽管能通过显式置信度评分识别说话人不可靠,却未在行为上加以利用。反而默认采取低效策略,如不顾伙伴的过度追问和在不确定时拒绝提问而强行猜测。BWIM为评估交互场景下的在线伙伴适应与语境推理提供了受控环境。
原文摘要 · Abstract (English)
We investigate the separation of literal interpretation from contextual inference in a collaborative block-building tasks, where an agent must resolve underspecified instructions using context. We adapt an existing two-speaker psycholinguistic paradigm into an interactive benchmark called Build What I Mean (BWIM). This setup contrasts a pragmatically cooperative speaker with one who is only literally reliable. In BWIM, models face underspecified instructions and must choose between making a contextual inference or requesting clarification at a small communication cost. Evaluating several state-of-the-art LLMs, we find a clear dissociation between judgment and action. Although models successfully detect speaker unreliability in explicit confidence ratings, they fail to leverage this awareness when taking action. Instead of deploying efficient clarification strategies, models default to suboptimal behaviors. These include partner-blind over-clarification and question-averse guessing under uncertainty. BWIM provides a controlled environment to evaluate online partner adaptation and contextual reasoning in interactive settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。