用人体感知模式构建智能体的推理基础,让机器像人一样理解世界。
Grounding Agent Reasoning in Image Schemas: A Neurosymbolic Approach to Embodied Cognition
- 将人类感官经验模式转化为形式化概念结构,用于指导智能体推理。
- 通过大模型将自然语言转为感知模式表示,实现神经与符号结合。
- 提升系统可解释性与人机交互直观性,适合具身智能研究者。
尽管具身人工智能取得进展,现有智能体推理系统仍难以捕捉人类理解与互动环境时所依赖的基本概念结构。为此,我们提出一种新框架,将具身认知理论与智能体系统结合,利用图像图式(image schemas)的形式化表征——即由感官运动体验构成的重复性模式,这些模式塑造了人类认知。通过定制大语言模型,将自然语言描述转换为基于这些感官运动模式的正式表达,构建一个神经符号系统,使智能体的理解扎根于基本概念结构。该方法不仅能提升效率与可解释性,还能通过共享的具身理解实现更直观的人机交互。
原文摘要 · Abstract (English)
Despite advances in embodied AI, agent reasoning systems still struggle to capture the fundamental conceptual structures that humans naturally use to understand and interact with their environment. To address this, we propose a novel framework that bridges embodied cognition theory and agent systems by leveraging a formal characterization of image schemas, which are defined as recurring patterns of sensorimotor experience that structure human cognition. By customizing LLMs to translate natural language descriptions into formal representations based on these sensorimotor patterns, we will be able to create a neurosymbolic system that grounds the agent's understanding in fundamental conceptual structures. We argue that such an approach enhances both efficiency and interpretability while enabling more intuitive human-agent interactions through shared embodied understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。