arXiv:2509.08421cs.CVcs.AI2025-09被引 1

解决多视角行人追踪中特征融合不均问题,提升追踪精度与鲁棒性。

ATLASFusion: Aggregation Tracking with Location-Aware Sparse Fusion for Robust Spatio-Temporal Multi-View Pedestrian Tracking

  • 仅投影有效特征点,避免插值失真
  • 密度加权聚合,提升可靠区域可信度
  • 各视角独立监督,增强跨视角一致性

为实现时空多媒体智能,多视角多目标追踪(MVMOT)面临跨摄像头保持目标身份一致性的挑战,导致追踪不准。现代方法的误差主要源于将多视角特征投影至统一鸟瞰图(BEV)空间时的特征表示扭曲。该投影常导致特征密度不均,降低融合表示的可靠性。为此,本文提出ATLASFusion,一种基于三项互补技术的可靠性感知稀疏BEV融合框架:稀疏透视变换仅投影有效特征点以避免插值伪影;密度感知加权聚合为空间可靠区域分配更高置信度;每视角BEV监督在融合前独立监督各相机的BEV特征。在WildTrack和MultiViewX基准上的实验表明,相比基线,ATLASFusion在IDF1、MODP和鲁棒性上均有提升。在WildTrack上取得95.9% IDF1,为对比方法中最高;在MultiViewX上定位精度(MODP)从75.0%提升至89.2%。在标定噪声下表现优异:当注入高幅噪声时,TrackTacular MODA降至0.0%,而ATLASFusion仍保持40.1% MODA;输入分辨率减半时,MVTr完全失效,TrackTacular MODA下降2.7%,而ATLASFusion仅下降1.1%至92.5% MODA。这些性能提升几乎无额外计算开销。

原文摘要 · Abstract (English)

For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining consistent object identities across different camera views, leading to tracking inaccuracies. A key source of these errors in modern methods is the distortion of feature representations when projecting from multiple views into a unified Bird's-Eye-View (BEV) space. This projection often creates non-uniform feature densities, harming the reliability of the fused representation. To address this, we present ATLASFusion, a reliability-aware sparse BEV fusion framework built on three complementary techniques: a sparse perspective transform that projects only valid feature points to prevent interpolation artifacts, density-aware weighted aggregation that assigns higher confidence to spatially reliable regions, and per-view BEV supervision that supervises each camera's BEV features independently before fusion. Experiments on the WildTrack and MultiViewX benchmarks demonstrated improvements in IDF1, MODP, and robustness over the baseline. ATLASFusion achieved a 95.9\% IDF1 tracking score on WildTrack, the highest among the compared methods and improved the localization precision (MODP) from 75.0\% to 89.2\% on MultiViewX. ATLASFusion also exhibited strong robustness under calibration noise. When high-magnitude noise was injected, TrackTacular failed entirely at 0.0\% MODA, whereas ATLASFusion retained 40.1\% MODA. When the input resolution was reduced by half, MVTr failed entirely and TrackTacular degraded by 2.7\% MODA, while ATLASFusion sustained 92.5\% MODA with only a 1.1\% decrease. These gains were achieved with negligible computational overhead.

多视角追踪鸟瞰图融合行人追踪稀疏建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。