用大模型提升车载AI通信效率,关键信息传输更精准。
Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks
- 基于大视觉语言模型,只传用户关注区域的图像片段。
- 在低信噪比下准确率提升超33%,12dB时提升13.4%。
- 适合车联网中资源受限、需快速响应的智能交互场景。
面向任务的语义通信已成为提升各类通信场景性能的关键方法。尽管生成式人工智能(如大语言模型)已被用于语义通信设计,但大模态模型(LMM)的潜力尚未充分挖掘。本文基于大语言与视觉助手(LLaVA)构建了车载AI助手,并提出一种面向任务的语义通信框架,以实现用户与云端服务器间的高效交互。为降低计算开销并缩短响应时间,我们优化了LLaVA的图像切片策略,仅聚焦用户最关注的区域。同时,结合客观与主观用户注意力评估图像块重要性,动态调整传输能耗,提升资源利用效率,确保关键信息精确传递。我们构建了一个面向交通场景的视觉问答(VQA)数据集进行评估。实验结果表明,在相同信道条件下,该框架显著提升了问题回答准确率,尤其在低信噪比环境下表现突出:在12dB SNR下准确率提升13.4%,在10dB SNR下提升33.1%。
原文摘要 · Abstract (English)
Task-oriented semantic communication has emerged as a fundamental approach for enhancing performance in various communication scenarios. While recent advances in Generative Artificial Intelligence (GenAI), such as Large Language Models (LLMs), have been applied to semantic communication designs, the potential of Large Multimodal Models (LMMs) remains largely unexplored. In this paper, we investigate an LMM-based vehicle AI assistant using a Large Language and Vision Assistant (LLaVA) and propose a task-oriented semantic communication framework to facilitate efficient interaction between users and cloud servers. To reduce computational demands and shorten response time, we optimize LLaVA's image slicing to selectively focus on areas of utmost interest to users. Additionally, we assess the importance of image patches by combining objective and subjective user attention, adjusting energy usage for transmitting semantic information. This strategy optimizes resource utilization, ensuring precise transmission of critical information. We construct a Visual Question Answering (VQA) dataset for traffic scenarios to evaluate effectiveness. Experimental results show that our semantic communication framework significantly increases accuracy in answering questions under the same channel conditions, performing particularly well in environments with poor Signal-to-Noise Ratios (SNR). Accuracy can be improved by 13.4% at an SNR of 12dB and 33.1% at 10dB, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。