用多模态大模型提升物体定位地图的语义准确性,减少误匹配。
Semantic Enhancement for Object SLAM with Heterogeneous Multimodal Large Language Model Agents
- 引入异构多模态大模型代理,动态增强场景语义理解。
- 通过代价矩阵优化数据关联,使语义准确率提升且误报减少。
- 异步处理机制大幅提速,适合实时机器人应用。
物体同时定位与建图(Object SLAM)系统在密集室内环境或场景变化时,难以正确关联语义相似的物体。本文提出SEO-SLAM框架,通过集成异构多模态大语言模型(MLLM)代理,增强语义建图能力并实现场景自适应。为提高计算效率,设计了异步处理方案,在不损失语义精度或SLAM性能的前提下显著降低推理时间。此外,提出结合语义距离与马氏距离的多数据关联策略,将问题建模为线性分配问题(LAP),有效缓解感知混淆。实验表明,SEO-SLAM在语义准确率和误报率方面优于基线方法;异步MLLM代理相比同步设置显著提升处理效率。还验证了其对下游任务如机器人辅助的改进潜力。数据集已公开:jungseokhong.com/SEO-SLAM。
原文摘要 · Abstract (English)
Object Simultaneous Localization and Mapping (SLAM) systems struggle to correctly associate semantically similar objects in close proximity, especially in cluttered indoor environments and when scenes change. We present Semantic Enhancement for Object SLAM (SEO-SLAM), a novel framework that enhances semantic mapping by integrating heterogeneous multimodal large language model (MLLM) agents. Our method enables scene adaptation while maintaining a semantically rich map. To improve computational efficiency, we propose an asynchronous processing scheme that significantly reduces the agents' inference time without compromising semantic accuracy or SLAM performance. Additionally, we introduce a multi-data association strategy using a cost matrix that combines semantic and Mahalanobis distances, formulating the problem as a Linear Assignment Problem (LAP) to alleviate perceptual aliasing. Experimental results demonstrate that SEO-SLAM consistently achieves higher semantic accuracy and reduces false positives compared to baselines, while our asynchronous MLLM agents significantly improve processing efficiency over synchronous setups. We also demonstrate that SEO-SLAM has the potential to improve downstream tasks such as robotic assistance. Our dataset is publicly available at: jungseokhong.com/SEO-SLAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。