6G时代智能图像传输,用AI提取关键信息降负载保画质
Semantic-Aware Visual Information Transmission With Key Information Extraction Over Wireless Networks
- 用AI动态提取人物姿态与背景,分离核心视觉信息
- 低信噪比下峰值信噪比显著优于传统方法
- 适合资源受限的移动多媒体应用,如远程医疗
6G网络要求前所未有的智能性、适应性和效率,以应对动态环境中超高速传输、超低延迟和海量连接的挑战。传统无线图像传输框架依赖静态配置和孤立的源信道编码,在波动信道条件下难以兼顾计算效率、鲁棒性与质量。为此,本文提出一种面向资源受限6G网络的AI原生深度联合源信道编码(JSCC)框架。通过集成关键信息提取与自适应背景合成,实现智能语义感知传输。利用AI工具Mediapipe进行人体姿态检测,Rembg实现背景移除,模型动态分离前景特征并从预训练库中匹配背景,降低数据负载的同时保持视觉保真度。实验表明,该方法在低信噪比(low-SNR)条件下相比传统JSCC方法显著提升峰值信噪比(PSNR),为资源受限的移动通信中的多媒体服务提供了可行解决方案。
原文摘要 · Abstract (English)
The advent of 6G networks demands unprecedented levels of intelligence, adaptability, and efficiency to address challenges such as ultra-high-speed data transmission, ultra-low latency, and massive connectivity in dynamic environments. Traditional wireless image transmission frameworks, reliant on static configurations and isolated source-channel coding, struggle to balance computational efficiency, robustness, and quality under fluctuating channel conditions. To bridge this gap, this paper proposes an AI-native deep joint source-channel coding (JSCC) framework tailored for resource-constrained 6G networks. Our approach integrates key information extraction and adaptive background synthesis to enable intelligent, semantic-aware transmission. Leveraging AI-driven tools, Mediapipe for human pose detection and Rembg for background removal, the model dynamically isolates foreground features and matches backgrounds from a pre-trained library, reducing data payloads while preserving visual fidelity. Experimental results demonstrate significant improvements in peak signal-to-noise ratio (PSNR) compared with traditional JSCC method, especially under low-SNR conditions. This approach offers a practical solution for multimedia services in resource-constrained mobile communications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。