arXiv:2609.08346cs.CVcs.AI2026-09

融合雷达与多模态视觉,实现高鲁棒性运动目标分割与追踪

Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking

论文配图:Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking
图 1 · 摘自论文原文
  • 基于雷达直接测量径向速度,结合图像特征提升运动检测精度
  • 在低光照、遮挡等恶劣条件下仍保持目标身份一致性,追踪准确率显著提升
  • 适合复杂环境下的安防监控系统,尤其适用于夜间或恶劣天气场景

运动目标感知需判断图像中哪些区域对应真实运动,并在时间上持续识别每个目标。依赖外观、光流或轨迹估计的方法在弱光、恶劣天气、反射和遮挡下表现不佳。雷达可直接测量径向速度,是天然解决方案。然而现有基准缺乏同步的雷达数据、密集运动实例掩码及时间一致的身份标签。为此,我们构建了RGBTR-Motion,一个同步校准的固定摄像头基准,包含RGB、热成像与雷达流,配有密集实例掩码和跨场景一致的身份信息。同时提出SAM-Radar框架,基于SAM3架构,融合校准后的RGBT特征与投影至图像位置的雷达回波。通过雷达回波的前景分类监督,无需文本提示即可有效剔除杂波。追踪器将有效雷达回波关联至轨迹,作为物理证据,在短暂视觉缺失或遮挡后仍能重建目标身份,避免重新初始化。SAM-Radar取得0.7027 IoU、0.8090 F1-50,MOTA、HOTA、IDF1分别提升0.2977、0.1603、0.2857。

原文摘要 · Abstract (English)

Moving-object perception must decide which image regions correspond to real motion and keep every instance identified over time. Methods that read motion from appearance, optical flow, or estimated trajectories lose that evidence under poor illumination, adverse weather, reflections, and occlusion. Radar is a natural remedy because it measures radial velocity directly instead of inferring it from photometric correspondence. However, existing benchmarks do not jointly provide radar measurements, dense moving-instance masks, and temporally consistent identities for surveillance. We therefore introduce RGBTR-Motion, a synchronized and calibrated fixed-camera benchmark that pairs RGB, thermal, and radar streams with dense instance masks and temporally consistent identities across diverse surveillance scenes. We also develop SAM-Radar, an RGB, thermal, and radar-based segmentation and tracking framework built on SAM 3. SAM-Radar's radar-aware detector fuses calibrated RGBT features with radar returns that are grounded at their projected image locations, and motion supervision, implemented as foreground classification of those projected returns, teaches the detector to reject clutter without any text prompt. The tracker associates accepted radar returns with individual trajectories and uses them as physical evidence that a visually degraded target remains present. This allows it to bridge short periods of low visibility or occlusion and reconnect a reappearing target to its existing identity instead of starting a new track. SAM-Radar attains 0.7027 IoU and 0.8090 F1-50, and raises MOTA, HOTA, and IDF1 by 0.2977, 0.1603, and 0.2857 over the strongest competing values.

运动分割雷达感知多模态目标追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。