用检索增强生成技术提升大模型对无线环境的多模态感知能力
ENWAR: A RAG-empowered Multi-Modal LLM Framework for Wireless Environment Perception
- 融合雷达、摄像头等多模态数据,通过检索增强实现环境认知
- 在DeepSense6G数据集上达到86%忠实度、80%正确性,显著优于通用模型
- 适合6G网络管理、智能交通等需实时环境理解的场景
大型语言模型(LLMs)在6G及未来网络的网络管理和编排中具有巨大潜力。然而,现有模型在特定领域知识和处理多模态传感数据方面存在局限,而这对于动态无线环境中的实时态势感知至关重要。本文提出ENWAR——一种环境感知增强的检索增强生成多模态大模型框架。ENWAR可无缝整合多模态感知输入,实现对复杂无线环境的感知、解读与认知处理,提供人类可理解的态势感知信息。在DeepSense6G数据集的GPS、LiDAR与摄像头模态组合上,结合Mistral-7b/8x7b和LLaMa3.1-8/70/405b等先进LLM进行评估。相较于这些基础模型常给出的泛化描述,ENWAR实现了更丰富的空间分析,准确识别位置,分析障碍物,并评估车辆间视距情况。结果表明,ENWAR在关键性能指标上最高达70%相关性、55%上下文召回率、80%正确性和86%忠实度,证明其在多模态感知与解释方面的有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) hold significant promise in advancing network management and orchestration in 6G and beyond networks. However, existing LLMs are limited in domain-specific knowledge and their ability to handle multi-modal sensory data, which is critical for real-time situational awareness in dynamic wireless environments. This paper addresses this gap by introducing ENWAR, an ENvironment-aWARe retrieval augmented generation-empowered multi-modal LLM framework. ENWAR seamlessly integrates multi-modal sensory inputs to perceive, interpret, and cognitively process complex wireless environments to provide human-interpretable situational awareness. ENWAR is evaluated on the GPS, LiDAR, and camera modality combinations of DeepSense6G dataset with state-of-the-art LLMs such as Mistral-7b/8x7b and LLaMa3.1-8/70/405b. Compared to general and often superficial environmental descriptions of these vanilla LLMs, ENWAR delivers richer spatial analysis, accurately identifies positions, analyzes obstacles, and assesses line-of-sight between vehicles. Results show that ENWAR achieves key performance indicators of up to 70% relevancy, 55% context recall, 80% correctness, and 86% faithfulness, demonstrating its efficacy in multi-modal perception and interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。