首个大规模水下声学目标追踪基准与高效追踪框架
SonarT165: A Large-scale Benchmark and STFTrack Framework for Acoustic Object Tracking
- 提出多视角模板融合与轨迹优化双模块,提升声学图像追踪精度
- 在165组方阵、165组扇形序列上实现领先性能,标注超20万条
- 适合水下机器人、海洋监测等场景的视觉追踪研究者使用
水下观测系统通常结合光学相机与成像声呐。当水下能见度不足时,仅声呐系统可提供稳定数据,因此亟需开展水下声学目标追踪(UAOT)研究。以往研究多采用传统方法和孪生网络,但缺乏统一评估基准,严重制约了方法价值。为此,本文提出首个大规模UAOT基准SonarT165,包含165组方阵序列、165组扇形序列及20.5万条高质量标注。实验表明,该基准揭示了现有SOT追踪器的局限性。为解决此问题,我们提出STFTrack框架,包含两个新模块:多视图模板融合模块(MTFM)通过交叉注意力机制融合原图与动态模板二值图的时空特征;最优轨迹修正模块(OTCM)引入声响应等效像素特性,设计归一化像素亮度响应分数,抑制由卡尔曼滤波预测框不准导致的次优匹配。此外,还引入声学图像增强方法与频率增强模块(FEM)进一步优化特征表示。大量实验验证,STFTrack在所提基准上达到当前最佳性能。代码已开源:https://github.com/LiYunfengLYF/SonarT165。
原文摘要 · Abstract (English)
Underwater observation systems typically integrate optical cameras and imaging sonar systems. When underwater visibility is insufficient, only sonar systems can provide stable data, which necessitates exploration of the underwater acoustic object tracking (UAOT) task. Previous studies have explored traditional methods and Siamese networks for UAOT. However, the absence of a unified evaluation benchmark has significantly constrained the value of these methods. To alleviate this limitation, we propose the first large-scale UAOT benchmark, SonarT165, comprising 165 square sequences, 165 fan sequences, and 205K high-quality annotations. Experimental results demonstrate that SonarT165 reveals limitations in current state-of-the-art SOT trackers. To address these limitations, we propose STFTrack, an efficient framework for acoustic object tracking. It includes two novel modules, a multi-view template fusion module (MTFM) and an optimal trajectory correction module (OTCM). The MTFM module integrates multi-view feature of both the original image and the binary image of the dynamic template, and introduces a cross-attention-like layer to fuse the spatio-temporal target representations. The OTCM module introduces the acoustic-response-equivalent pixel property and proposes normalized pixel brightness response scores, thereby suppressing suboptimal matches caused by inaccurate Kalman filter prediction boxes. To further improve the model feature, STFTrack introduces a acoustic image enhancement method and a Frequency Enhancement Module (FEM) into its tracking pipeline. Comprehensive experiments show the proposed STFTrack achieves state-of-the-art performance on the proposed benchmark. The code is available at https://github.com/LiYunfengLYF/SonarT165.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。