arXiv:2608.23249cs.CVeess.SP2026-08

用多对天线融合提升射频成像的语义重建与3D检测,抗噪能力更强。

Semantic Reconstruction and 3-D Detection via Learned Multi-Pair Fusion in RF Imaging

论文配图:Semantic Reconstruction and 3-D Detection via Learned Multi-Pair Fusion in RF Imaging
图 1 · 摘自论文原文
  • 通过训练3D U-Net融合六对天线的图像,实现隐式融合与显式分类
  • 在低信噪比下仍保持稳定检测,远优于传统强度重建方法
  • 引入未知类机制,可正确识别未训练的新物体而非误标

针对存在各向异性的多站射频成像问题,反射信号依赖于发射(Tx)和接收(Rx)阵列位置。目标是将视场内体素标记为有限语义类别并聚类为物体实例。对每对Tx--Rx采用标准逆问题求解器生成图像,再输入训练好的三维U-Net进行隐式融合与显式体素分类。在受控、欠定的多站设置下,比较单快照的后向投影(BP)与LASSO,以及多快照的非相干BP与组LASSO成像方法。每种方法训练独立的U-Net,融合六对Tx--Rx输入,输出每个体素的类别概率向量。取最大概率类别得到标签体数据——语义重建;再通过几何后处理(聚类与主成分分析)获得物体实例及其方向包围盒。在宽范围信噪比下,语义重建(以分割交并比评分)与3D检测性能衰减远比经典强度重建平缓,尤其检测在强度图完全消失的噪声水平下仍可靠。因真实场景含未训练类别物体,引入显式未知类,通过异常暴露训练,将未见物体标记为未知,避免误标为已知类别。

原文摘要 · Abstract (English)

We consider a multistatic radio-frequency imaging problem with anisotropy, in which the reflection from a point depends on the positions of the transmit (Tx) and receive (Rx) arrays. The goal is to label the voxels of a field of view by a finite set of semantic classes and to group them into object instances. For the image formation of each Tx--Rx pair we apply a standard inverse-problem solver, and we feed the resulting per-pair reconstructions into a trained three-dimensional (3-D) U-Net that performs the fusion implicitly and the per-voxel classification explicitly. On a controlled, under-determined multistatic setup, we consider the following image formation methods: back-projection (BP) and the least absolute shrinkage and selection operator (LASSO) from a single deterministic snapshot, and incoherent BP and group-LASSO from multiple fading snapshots. For each imaging method we train a separate U-Net that fuses the six Tx--Rx pairs (its input channels) and assigns each voxel a probability vector over the classes. Taking the most probable class gives a labeled volume---the semantic reconstruction. Object instances and their oriented bounding boxes then follow by geometric post-processing (clustering and principal-component analysis). Across a wide range of signal-to-noise ratio, the semantic reconstruction (scored against ground truth by segmentation intersection-over-union) and the resulting 3-D detection degrade far more gracefully than the classical intensity reconstruction: the detection in particular stays reliable well into noise levels at which that reconstruction has dissolved. Because real scenes contain objects of classes the network was not trained on, we add an explicit unknown class trained by outlier exposure, which labels held-out novel objects as unknown instead of mislabeling them as a known class by reconstructed shape.

射频成像3D检测语义分割抗噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。