通过频域建模传感器差异,提升雷达相机3D感知跨数据集泛化能力
Understanding Cross-Sensor Feature Variations for Generalizable 3D Perception

- 在频域建模视觉场景变化,合成多样化源域视图
- 发现图像变化对多模态BEV特征的影响模式,提升融合稳定性
- 训练时增强鲁棒性,推理无需修改,适合无目标域数据场景
雷达-相机鸟瞰图(BEV)感知在跨数据集评估时性能常下降,因驾驶场景、传感器配置和环境条件变化会改变输入观测与内部融合表征。本文从源域变化建模角度出发,旨在不依赖目标域样本的情况下提升BEV检测器的鲁棒性。提出一种框架,通过频域刻画视觉场景变化,并用于合成多样化的源域视图;通过对比生成的融合BEV表征,进一步捕捉图像级变化对多模态BEV特征的影响。这些变化模式被用于正则化检测器,促使学习到的融合空间在潜在场景变化下保持稳定。该方法仅在训练阶段使用,不影响推理流程。在View-of-Delft与TJ4DRadSet之间的跨数据集雷达-相机3D检测任务中,对多个BEV融合骨干网络均实现一致提升,且在少量目标域数据情况下仍有效。
原文摘要 · Abstract (English)
Radar-camera BEV perception often suffers from degraded performance when evaluated across datasets, as changes in driving scenes, sensor configurations, and environmental conditions can alter both the input observations and the internal fused representations. This work studies this issue from the perspective of source-domain variation modeling, aiming to improve the robustness of BEV-based 3D detectors without relying on target-domain samples. We introduce a framework that characterizes visual scene variations in the frequency domain and uses them to synthesize diverse source-domain views. By comparing the resulting fused BEV representations, the framework further captures how image-level variations influence multi-modal BEV features. These variation patterns are then used to regularize the detector, encouraging the learned fusion space to remain stable under latent scene changes. The proposed method is applied only during training and leaves the inference pipeline unchanged. Experiments on cross-dataset radar-camera 3D detection between View-of-Delft and TJ4DRadSet demonstrate consistent improvements over multiple BEV fusion backbones, and the gains remain effective when a small amount of target-domain data is available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。