arXiv:2503.07670cs.NIeess.IV2025-03中稿 · @ ICC 2025被引 10

用多模态检索增强生成,提升无线环境感知的智能优化效果。

Retrieval Augmented Generation with Multi-Modal LLM Framework for Wireless Environments

  • 通过多传感器数据融合与提示工程构建统一向量库
  • 相比传统方法在相关性、准确率等指标上提升7%-12%
  • 适合需要实时响应的无线网络智能优化场景

未来无线网络需在高数据速率、低功耗和无缝连接之间取得平衡,亟需强大优化能力。大型语言模型(LLMs)已用于通用优化场景。为利用生成式AI(GAI)模型,本文提出面向多传感器无线环境感知的检索增强生成(RAG)框架。通过领域特定提示工程,有效整合无线环境中多源异构传感器数据。提出关键预处理流程,包括图像转文本、目标检测与距离计算,以构建统一向量数据库,支撑全局无线任务中LLM的优化。在OpenAI GPT与Google Gemini模型上的评估表明,相较于传统基于LLM的设计,本方法在相关性、忠实性、完整性、相似性和准确性上分别提升8%、8%、10%、7%和12%。此外,基于向量数据库的RAG框架具备计算高效性,在延迟约束下实现实时收敛。

原文摘要 · Abstract (English)

Future wireless networks aim to deliver high data rates and lower power consumption while ensuring seamless connectivity, necessitating robust optimization. Large language models (LLMs) have been deployed for generalized optimization scenarios. To take advantage of generative AI (GAI) models, we propose retrieval augmented generation (RAG) for multi-sensor wireless environment perception. Utilizing domain-specific prompt engineering, we apply RAG to efficiently harness multimodal data inputs from sensors in a wireless environment. Key pre-processing pipelines including image-to-text conversion, object detection, and distance calculations for multimodal RAG input from multi-sensor data are proposed to obtain a unified vector database crucial for optimizing LLMs in global wireless tasks. Our evaluation, conducted with OpenAI's GPT and Google's Gemini models, demonstrates an 8%, 8%, 10%, 7%, and 12% improvement in relevancy, faithfulness, completeness, similarity, and accuracy, respectively, compared to conventional LLM-based designs. Furthermore, our RAG-based LLM framework with vectorized databases is computationally efficient, providing real-time convergence under latency constraints.

无线感知检索增强多模态LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。