让机器人在8GB内存设备上实现高精度自主决策
RDMM: Fine-Tuned LLM Models for On-Device Robotic Decision Making with Enhanced Contextual Awareness in Specific Domains
- 用领域专用模型增强机器人上下文理解能力
- 实测决策准确率达93%,可在8GB内存设备运行
- 支持视觉感知与实时语音交互,适合家庭服务场景
大型语言模型(LLMs)在将物理机器人与智能系统融合方面取得显著进展。本研究展示了框架在真实家庭竞赛环境中的能力。提出RDMM(机器人决策模型)框架,具备领域特定情境下的决策能力及对自身知识和能力的认知。该框架利用信息提升系统自主决策水平。相比其他方法,本工作聚焦实时、本地化解决方案,成功在仅8GB内存的硬件上运行。框架集成视觉感知模型,使机器人具备环境理解能力;同时引入实时语音识别,改善人机交互体验。实验结果表明,RDMM框架规划准确率达到93%。此外,本文构建了一个包含27,000个规划实例的新数据集,以及1,300个从竞赛中提取的图文标注样本。本工作开发的框架、基准、数据集和模型均已公开于GitHub:https://github.com/shadynasrat/RDMM。
原文摘要 · Abstract (English)
Large language models (LLMs) represent a significant advancement in integrating physical robots with AI-driven systems. We showcase the capabilities of our framework within the context of the real-world household competition. This research introduces a framework that utilizes RDMM (Robotics Decision-Making Models), which possess the capacity for decision-making within domain-specific contexts, as well as an awareness of their personal knowledge and capabilities. The framework leverages information to enhance the autonomous decision-making of the system. In contrast to other approaches, our focus is on real-time, on-device solutions, successfully operating on hardware with as little as 8GB of memory. Our framework incorporates visual perception models equipping robots with understanding of their environment. Additionally, the framework has integrated real-time speech recognition capabilities, thus enhancing the human-robot interaction experience. Experimental results demonstrate that the RDMM framework can plan with an 93\% accuracy. Furthermore, we introduce a new dataset consisting of 27k planning instances, as well as 1.3k text-image annotated samples derived from the competition. The framework, benchmarks, datasets, and models developed in this work are publicly available on our GitHub repository at https://github.com/shadynasrat/RDMM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。