用视觉信息辅助毫米波信道建模,大幅减少测量需求。
Taming Vision Priors for Data Efficient mmWave Channel Modeling
- 用摄像头图像提取语义特征,生成材料参数初始估计。
- 仅需几十次稀疏测量即可快速校准,误差降低59%。
- 适合需要快速部署的自动驾驶与增强现实场景。
精确建模毫米波(mmWave)传播对实时增强现实和自动驾驶系统至关重要。可微分射线追踪提供物理基础解法,但因过度依赖全面信道测量或脆弱的手动调参场景模型而难以部署。我们提出VisRFTwin,一种可扩展且数据高效的数字孪生框架,将视觉导出的材料先验与可微分射线追踪结合。利用普通摄像头的多视角图像,通过冻结的视觉-语言模型提取密集语义嵌入,并转化为场景表面介电常数和电导率的初始估计。这些先验用于初始化基于Sionna的可微分射线追踪器,仅需数十次稀疏信道探测即可通过梯度下降快速校准材料参数。校准后,视觉特征与材料参数的关联被保留,使新场景可快速迁移而无需重复校准。在办公室内部、城市峡谷和动态公共空间三个真实场景中评估显示,VisRFTwin将信道测量需求减少最多达10倍,同时比纯数据驱动深度学习方法降低59%的中位延迟扩展误差。
原文摘要 · Abstract (English)
Accurately modeling millimeter-wave (mmWave) propagation is essential for real-time AR and autonomous systems. Differentiable ray tracing offers a physics-grounded solution but still facing deployment challenges due to its over-reliance on exhaustive channel measurements or brittle, hand-tuned scene models for material properties. We present VisRFTwin, a scalable and data-efficient digital-twin framework that integrates vision-derived material priors with differentiable ray tracing. Multi-view images from commodity cameras are processed by a frozen Vision-Language Model to extract dense semantic embeddings, which are translated into initial estimates of permittivity and conductivity for scene surfaces. These priors initialize a Sionna-based differentiable ray tracer, which rapidly calibrates material parameters via gradient descent with only a few dozen sparse channel soundings. Once calibrated, the association between vision features and material parameters is retained, enabling fast transfer to new scenarios without repeated calibration. Evaluations across three real-world scenarios, including office interiors, urban canyons, and dynamic public spaces show that VisRFTwin reduces channel measurement needs by up to 10$\times$ while achieving a 59% lower median delay spread error than pure data-driven deep learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。