arXiv:2510.01531cs.AIcs.CL2025-10

让AI主动查信息,提升在模糊环境下的决策能力

Information Seeking for Robust Decision Making under Partial Observability

  • 用大模型主动规划探查动作,校准内部认知与真实环境
  • 在部分可观测场景中性能比旧方法高74%,且不增加样本消耗
  • 适用于机器人操作、网页导航等任务,适配多种大模型

在信息不全、动态噪声大的实际环境中,人类解决问题时需主动获取信息以更新内部认知并指导后续决策。当真实环境状态不可直接观测时,主动信息获取对决策至关重要。尽管现有大语言模型(LLM)规划代理已处理观测不确定性,但常忽略其内部动态与真实环境间的差异。本文提出信息寻求决策规划器(InfoSeeker),一个将任务导向规划与信息获取相结合的LLM决策框架,旨在部分可观测环境下通过对齐内部动态实现最优决策。InfoSeeker引导大模型主动规划探查动作,验证理解、检测环境变化或测试假设,再生成或修正任务计划。为评估该方法,我们构建了一个包含不完整观测和不确定动态的新基准套件。实验表明,InfoSeeker相比之前方法取得74%的绝对性能提升,且不牺牲样本效率。此外,InfoSeeker在不同大模型间具有泛化能力,并在机器人操作与网页导航等经典基准上超越基线。这些结果凸显了将规划与信息获取紧密结合对提升部分可观测环境下鲁棒行为的重要性。

原文摘要 · Abstract (English)

Explicit information seeking is essential to human problem-solving in practical environments characterized by incomplete information and noisy dynamics. When the true environmental state is not directly observable, humans seek information to update their internal dynamics and inform future decision-making. Although existing Large Language Model (LLM) planning agents have addressed observational uncertainty, they often overlook discrepancies between their internal dynamics and the actual environment. We introduce Information Seeking Decision Planner (InfoSeeker), an LLM decision-making framework that integrates task-oriented planning with information seeking to align internal dynamics and make optimal decisions under uncertainty in both agent observations and environmental dynamics. InfoSeeker prompts an LLM to actively gather information by planning actions to validate its understanding, detect environmental changes, or test hypotheses before generating or revising task-oriented plans. To evaluate InfoSeeker, we introduce a novel benchmark suite featuring partially observable environments with incomplete observations and uncertain dynamics. Experiments demonstrate that InfoSeeker achieves a 74% absolute performance gain over prior methods without sacrificing sample efficiency. Moreover, InfoSeeker generalizes across LLMs and outperforms baselines on established benchmarks such as robotic manipulation and web navigation. These findings underscore the importance of tightly integrating planning and information seeking for robust behavior in partially observable environments. The project page is available at https://infoseekerllm.github.io

大模型决策信息获取部分可观测规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。