测试大模型在推理中主动获取信息的能力,发现其常过早下结论或无法有效验证假设。
Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

- 设计交互式游戏探测模型如何逐步获取证据并更新假设
- 提前给全证据时成功率更高,分步获取时模型易过早决策
- 模型自己提问效果差,但自选证据的假设更一致
溯因推理要求生成能解释观测证据的假设,并在新证据出现时进行修正。尽管大语言模型(LLMs)常被评估是否能正确解决溯因任务,但它们如何获取证据、更新假设以及决定何时停止仍不明确。我们引入了「外星绑架」游戏,一种用于研究不同交互模式下行为的互动探针。该模式在证据是否提前提供、以及查询由模型自主选择还是由代理(oracle)提供之间变化。结果表明,提前提供全部证据时模型成功率更高;在多轮交互中,部分模型过早形成假设,另一些则耗尽回合数仍未收敛。当由代理提供示例时,模型表现优于自主提问,但其最终假设与所选证据更一致。这表明模型可能基于自选证据形成假设,却未能充分区分替代方案,且难以验证和精炼假设,或判断停止时机。
原文摘要 · Abstract (English)
Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language models (LLMs) are often evaluated on whether they solve abductive reasoning tasks correctly, less is known about how they acquire evidence, update their hypotheses, and decide when to stop. We introduce Alien Abduction game, an interactive probe for studying these behaviours under different interaction modes. The modes vary in whether evidence is provided upfront or across turns, and whether queries are selected by the model or examples are provided by the oracle. Across models, providing evidence upfront leads to higher success rates than distributing it across turns. In multi-turn settings, some models commit before using the available evidence, while others exhaust the turn budget without converging. Models also achieve higher success rates when examples are provided by the oracle than when they select their own queries, although their final hypotheses are more consistent with the evidence they selected. These findings suggest that models may form hypotheses that fit self-selected evidence without sufficiently distinguishing them from alternatives, and may struggle to validate and refine their hypotheses or determine when to stop.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。