用WiFi信号生成高清环境图像,还能通过文字控制生成内容。
High-resolution efficient image generation from WiFi CSI using a pretrained latent diffusion model
- 用轻量网络将WiFi信号直接映射到扩散模型潜空间
- 生成图像质量优于同类方法,计算效率更高
- 支持文字引导生成,适合需要可控图像的应用
我们提出LatentCSI,一种从WiFi信道状态信息(CSI)生成物理环境高分辨率图像的新方法,该方法基于预训练的潜空间扩散模型(LDM)。与依赖复杂计算的GAN等传统方法不同,本方法采用轻量级神经网络将CSI幅度直接映射至LDM的潜空间,随后在潜空间中使用文本引导的去噪扩散模型进行重构,并通过LDM的预训练解码器生成最终图像。该设计跳过了像素空间生成和显式图像编码阶段,显著提升生成效率与图像质量。我们在两个数据集上验证:自采的宽带CSI数据集(使用市售WiFi设备与摄像头同步采集)及公开的MM-Fi数据集子集。结果表明,LatentCSI在计算效率和感知质量上均优于同等复杂度的基线方法,且具备独特的文本可控性优势。
原文摘要 · Abstract (English)
We present LatentCSI, a novel method for generating images of the physical environment from WiFi CSI measurements that leverages a pretrained latent diffusion model (LDM). Unlike prior approaches that rely on complex and computationally intensive techniques such as GANs, our method employs a lightweight neural network to map CSI amplitudes directly into the latent space of an LDM. We then apply the LDM's denoising diffusion model to the latent representation with text-based guidance before decoding using the LDM's pretrained decoder to obtain a high-resolution image. This design bypasses the challenges of pixel-space image generation and avoids the explicit image encoding stage typically required in conventional image-to-image pipelines, enabling efficient and high-quality image synthesis. We validate our approach on two datasets: a wide-band CSI dataset we collected with off-the-shelf WiFi devices and cameras; and a subset of the publicly available MM-Fi dataset. The results demonstrate that LatentCSI outperforms baselines of comparable complexity trained directly on ground-truth images in both computational efficiency and perceptual quality, while additionally providing practical advantages through its unique capacity for text-guided controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。