让AI在生成3D场景时像人一样多源调用知识,提升多样性。
MoK-RAG: Mixture of Knowledge Paths Enhanced Retrieval-Augmented Generation for Embodied AI Environments
- 将大模型知识按功能拆分,支持多路径检索。
- 自动与人工评估均显示生成场景更丰富多样。
- 适合需要复杂环境生成的具身智能研究者。
人类在决策时会从多种专业知识源中提取信息,而现有检索增强生成(RAG)系统通常仅依赖单一知识源,造成认知与算法的差距。为此,我们提出MoK-RAG,一种多源RAG框架,通过将大语言模型(LLM)语料库功能分区,实现多条专业化知识路径的检索。应用于3D模拟环境生成时,MoK-RAG3D进一步将3D资源按层级知识树结构划分。不同于以往仅依赖人工评估的方法,我们首次引入自动化评估3D场景的方法。实验表明,无论是自动还是人工评估,MoK-RAG3D均能有效提升具身智能体生成场景的多样性。
原文摘要 · Abstract (English)
While human cognition inherently retrieves information from diverse and specialized knowledge sources during decision-making processes, current Retrieval-Augmented Generation (RAG) systems typically operate through single-source knowledge retrieval, leading to a cognitive-algorithmic discrepancy. To bridge this gap, we introduce MoK-RAG, a novel multi-source RAG framework that implements a mixture of knowledge paths enhanced retrieval mechanism through functional partitioning of a large language model (LLM) corpus into distinct sections, enabling retrieval from multiple specialized knowledge paths. Applied to the generation of 3D simulated environments, our proposed MoK-RAG3D enhances this paradigm by partitioning 3D assets into distinct sections and organizing them based on a hierarchical knowledge tree structure. Different from previous methods that only use manual evaluation, we pioneered the introduction of automated evaluation methods for 3D scenes. Both automatic and human evaluations in our experiments demonstrate that MoK-RAG3D can assist Embodied AI agents in generating diverse scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。