arXiv:2508.14564cs.AIcs.CL2025-08中稿 · ICSR25被引 1

用结构化推理路径提升大模型视角理解能力

Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs

  • 构建三类结构化示例:目标路径、信息节点路径、决策对比序列
  • 在基础注意力任务中表现良好,但对遮挡空间推理仍失败
  • 需显式信念追踪和成本建模才能实现真实社会协作

大语言模型在自主代理的视角理解能力方面面临挑战,尤其在主动感知、协作推理和立场理解(即理解另一代理能看见或知道什么)的任务上。本研究探索使用由Fast Downward规划器生成的转换解图构造的结构化示例,以提升基于ReAct框架的LLM代理性能。提出一种结构化解题流程,生成三类示例:最优目标路径(G型)、信息节点路径(E型)和对比不同行动步骤的最优决策序列(L型)。这些解通过提示大模型显式表达每一步推理,转化为“思考-动作”示例。尽管L型示例略微减少澄清请求和总动作步数,但未带来稳定改进。代理在需要基础注意力过滤的任务中成功,但在涉及遮挡空间心智模拟或权衡认知行为成本的情境中表现不佳。结果表明,仅靠结构化示例不足以实现稳健的视角理解,亟需显式信念追踪、成本建模及更丰富的环境支持,以实现基于社会情境的大模型协作。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) and reasoning frameworks have opened new possibilities for improving the perspective -taking capabilities of autonomous agents. However, tasks that involve active perception, collaborative reasoning, and perspective taking (understanding what another agent can see or knows) pose persistent challenges for current LLM-based systems. This study investigates the potential of structured examples derived from transformed solution graphs generated by the Fast Downward planner to improve the performance of LLM-based agents within a ReAct framework. We propose a structured solution-processing pipeline that generates three distinct categories of examples: optimal goal paths (G-type), informative node paths (E-type), and step-by-step optimal decision sequences contrasting alternative actions (L-type). These solutions are further converted into ``thought-action'' examples by prompting an LLM to explicitly articulate the reasoning behind each decision. While L-type examples slightly reduce clarification requests and overall action steps, they do not yield consistent improvements. Agents are successful in tasks requiring basic attentional filtering but struggle in scenarios that required mentalising about occluded spaces or weighing the costs of epistemic actions. These findings suggest that structured examples alone are insufficient for robust perspective-taking, underscoring the need for explicit belief tracking, cost modelling, and richer environments to enable socially grounded collaboration in LLM-based agents.

大模型推理视角理解认知建模智能体协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。