arXiv:2601.18242cs.CVcs.NI2026-01被引 3

用视觉语言模型指导光线追踪,快速准确估计多材料射频参数

Vision-Language-Model-Guided Differentiable Ray Tracing for Fast and Accurate Multi-Material RF Parameter Estimation

  • 用VLM解析图像生成材料先验,指导光线追踪初始化
  • 相比随机初始化,收敛速度提升2-4倍,误差降低10-100倍
  • 适合6G电磁数字孪生中少测量条件下的材料参数估计

准确的射频(RF)材料参数对6G系统的电磁数字孪生至关重要,但基于梯度的逆向光线追踪(RT)对初始值敏感且在测量有限时计算成本高。本文提出一种由视觉语言模型(VLM)引导的框架,加速并稳定多材料参数估计。VLM解析场景图像,推断材料类别,并通过ITU-R材料表映射为定量先验,提供导电性初始值;同时选择能激发多样、可区分材料路径的收发机位置。从这些先验出发,利用测量接收信号强度,在可微分光线追踪(DRT)引擎中进行梯度优化。NVIDIA Sionna中的室内场景实验表明,相比均匀或随机初始化及放置基线,收敛速度提升2-4倍,最终参数误差降低10-100倍,仅需少数接收机即可实现低于0.1%的平均相对误差。复杂度分析显示,每迭代时间近似线性随材料数和测量配置增长,而VLM引导的位置设计显著减少所需测量数。深度与光线条数的消融实验确认精度进一步提升,且每迭代开销无明显增加。结果证明,来自VLM的语义先验能有效引导物理驱动优化,实现快速可靠的射频材料估计。

原文摘要 · Abstract (English)

Accurate radio-frequency (RF) material parameters are essential for electromagnetic digital twins in 6G systems, yet gradient-based inverse ray tracing (RT) remains sensitive to initialization and costly under limited measurements. This paper proposes a vision-language-model (VLM) guided framework that accelerates and stabilizes multi-material parameter estimation in a differentiable RT (DRT) engine. A VLM parses scene images to infer material categories and maps them to quantitative priors via an ITU-R material table, yielding informed conductivity initializations. The VLM further selects informative transmitter/receiver placements that promote diverse, material-discriminative paths. Starting from these priors, the DRT performs gradient-based refinement using measured received signal strengths. Experiments in NVIDIA Sionna on indoor scenes show 2-4$\times$ faster convergence and 10-100$\times$ lower final parameter error compared with uniform or random initialization and random placement baselines, achieving sub-0.1\% mean relative error with only a few receivers. Complexity analyses indicate per-iteration time scales near-linearly with the number of materials and measurement setups, while VLM-guided placement reduces the measurements required for accurate recovery. Ablations over RT depth and ray counts confirm further accuracy gains without significant per-iteration overhead. Results demonstrate that semantic priors from VLMs effectively guide physics-based optimization for fast and reliable RF material estimation.

射频参数光线追踪视觉语言模型6G数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。