让语言模型像人一样理性提问和决策,提升信息探索效率。
Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- 设计对话任务模拟人类认知,评估并改进模型的提问与决策能力。
- 引入贝叶斯实验设计启发的推理策略,使模型提问更有效,信息增益提升94.2%。
- 低成本模型在任务中超越人类和顶级大模型,适合资源受限场景。
许多新兴AI应用(如科学发现、医疗诊断)需要智能体具备战略性信息获取能力:形成假设、提出针对性问题,并在不确定性下决策。在资源有限的高风险场景中,语言模型是否表现出理性代理行为?基于人类认知研究,我们开发了评估与增强代理信息获取的方法。首先,提出一种名为「协作战舰」的决策导向对话任务,其中指挥官需权衡探索(提问)与行动(开火),观察员需提供准确且上下文相关的回答。对比42名真人玩家,发现多数语言模型在提出有意义问题、生成准确答案及识别高价值行动方面表现不佳。为此,我们设计了受贝叶斯实验设计启发的蒙特卡洛推理策略。对于观察员模型,准确率最高提升14.7个百分点;对于指挥官模型,预期信息增益(EIG)最高提升0.227比特(达到理论噪声上限的94.2%)。结合两者,目标定位精度提升0.303-0.374 F1,使弱模型如Llama-4-Scout在成本仅为GPT-5的1%时,胜率从8%升至82%,超越人类;在与GPT-5对比中,胜率从0%升至67%。该方法在「猜谁?」任务中同样显著提升准确率(+28.3-42.4个百分点),证明其通用性。
原文摘要 · Abstract (English)
Many emerging applications of AI--from scientific discovery to medical diagnosis--require agents to seek information strategically: forming hypotheses, asking targeted questions, and making decisions under uncertainty. In high-stakes settings with limited resources, do language models (LMs) behave like rational agents? Drawing on insights from human cognition, we develop methods to evaluate and enhance agentic information-seeking. First, we introduce a decision-oriented dialogue task called Collaborative Battleship, in which a Captain must balance exploration (asking questions) and action (taking shots), while a Spotter must supply accurate, contextually-grounded answers. Compared to human players (N=42), we find that many LM agents struggle to ask informative questions, produce accurate answers, and identify high-utility actions. To address these gaps, we develop novel Monte Carlo inference strategies for LMs inspired by Bayesian Experimental Design (BED). For Spotter agents, our approach boosts accuracy by up to 14.7% absolute over LM-only baselines; for Captain agents, it raises expected information gain (EIG) by up to 0.227 bits (94.2% of the achievable noise ceiling). Combined, these components yield sharper targeting (+0.303-0.374 F1), and enable weaker LMs, such as Llama-4-Scout, to outperform both humans (8% -> 82% win rate) and frontier models (0% -> 67% win rate vs. GPT-5) at ~1% of GPT-5's cost. We replicate these findings on Guess Who?, where our methods significantly boost accuracy (+28.3-42.4 p.p.), demonstrating their general applicability for building information-seeking agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。