用文本大模型构建复杂物理世界,靠几何约束提升空间推理能力
Navigate Complex Physical Worlds via Geometrically Constrained LLM
- 通过多层图与多智能体框架,统一几何规则增强空间理解
- 利用遗传算法求解几何约束问题,实现多步多目标推理
- 适合对具身智能、物理模拟感兴趣的开发者与研究者
本研究探索仅基于文本知识的大型语言模型(LLMs)在重建与构建物理世界方面的潜力,并考察模型性能对空间理解能力的影响。为提升对复杂物理世界中几何与空间关系的理解,研究引入一套几何规范,构建基于多层图与多智能体系统框架的工作流。通过统一几何规范下的多层图,探究LLMs在空间环境中实现多步、多目标几何推理的机制。同时,借鉴大规模模型知识,采用遗传算法求解几何约束问题。研究表明,该方法创新性地验证了纯文本驱动的LLMs作为物理世界建造者的可行性,并设计出可扩展的能力增强工作流。
原文摘要 · Abstract (English)
This study investigates the potential of Large Language Models (LLMs) for reconstructing and constructing the physical world solely based on textual knowledge. It explores the impact of model performance on spatial understanding abilities. To enhance the comprehension of geometric and spatial relationships in the complex physical world, the study introduces a set of geometric conventions and develops a workflow based on multi-layer graphs and multi-agent system frameworks. It examines how LLMs achieve multi-step and multi-objective geometric inference in a spatial environment using multi-layer graphs under unified geometric conventions. Additionally, the study employs a genetic algorithm, inspired by large-scale model knowledge, to solve geometric constraint problems. In summary, this work innovatively explores the feasibility of using text-based LLMs as physical world builders and designs a workflow to enhance their capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。