用大模型生成符号网络,让机器人理解物品的上下文可用性。
Proposition of Affordance-Driven Environment Recognition Framework Using Symbol Networks in Large Language Models
- 通过大模型输出构建符号网络,提取隐含的物体使用逻辑。
- 以苹果为例,成功识别出不同场景下的可操作属性。
- 方法解释性强,适合需要常识推理的智能机器人应用。
为实现机器人与人类共存,理解动态情境并基于常识和可操作性选择适当动作至关重要。传统人工智能系统在应用可操作性时面临挑战,因其依赖源于常识的隐含知识。大型语言模型(LLMs)凭借其处理广泛人类知识的能力,提供了新机遇。本研究提出一种利用大模型输出自动获取可操作性的方法:首先生成文本,再通过形态学和依存分析重构为符号网络,并基于网络距离计算可操作性。以'apple'为例的实验表明,该方法能有效提取上下文相关的可操作性,且具有高可解释性。结果说明,由大模型输出重构的符号网络,使机器人能够高效理解可操作性,弥合了符号化数据与类人情境理解之间的差距。
原文摘要 · Abstract (English)
In the quest to enable robots to coexist with humans, understanding dynamic situations and selecting appropriate actions based on common sense and affordances are essential. Conventional AI systems face challenges in applying affordance, as it represents implicit knowledge derived from common sense. However, large language models (LLMs) offer new opportunities due to their ability to process extensive human knowledge. This study proposes a method for automatic affordance acquisition by leveraging LLM outputs. The process involves generating text using LLMs, reconstructing the output into a symbol network using morphological and dependency analysis, and calculating affordances based on network distances. Experiments using ``apple'' as an example demonstrated the method's ability to extract context-dependent affordances with high explainability. The results suggest that the proposed symbol network, reconstructed from LLM outputs, enables robots to interpret affordances effectively, bridging the gap between symbolized data and human-like situational understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。