提出分阶段空地协同感知框架,解决视角差异导致的干扰问题。
Rethinking Air-Ground Collaboration: A Progressive Cross-Task Benchmark and Socialized Learning Framework

- 构建分阶段协同任务,从全局定位逐步到目标关联与细粒度解析
- 在74.5万帧数据上实现平均下游性能提升7.86%,共进化增益3.73%
- 适合研究多智能体协同感知、跨视角视觉理解的开发者
空地协同感知对复杂动态环境中的鲁棒视觉理解至关重要。现有研究通常将协同视为单任务跨视角融合,忽略了定位、目标关联与细粒度解析之间的功能依赖。此外,空地视角的异质性带来显著的几何、尺度和遮挡差异,导致统一特征共享易引发负迁移。为此,我们把空地感知建模为分阶段跨任务协同任务,构建了包含超过74.5万帧原始视频的空地渐进式协同(AGPC)基准。基于该基准,提出社会化的协同感知(SCP)框架,从空中全局定位逐步推进至地面目标关联与身份感知。其核心模块双层路由器(DLR)解耦输入侧多尺度专家选择与输出侧任务条件调制,实现选择性跨视图与跨任务交互,抑制有害干扰。大量实验表明,SCP效果显著:取得3.73%的共进化增益和7.86%的平均下游性能提升,证明任务条件协同优于统一融合。代码已开源。
原文摘要 · Abstract (English)
Air-ground collaborative perception is crucial for robust visual understanding in real-world dynamic environments. However, existing studies typically formulate collaboration as single-task cross-view fusion, overlooking the functional dependencies among localization, target association, and fine-grained parsing. In addition, the heterogeneous nature of aerial and ground views introduces substantial geometric, scale, and occlusion discrepancies, making uniform feature sharing vulnerable to negative transfer. To tackle these issues, we model air-ground perception as a progressive cross-task collaboration task and construct the Air-Ground Progressive Collaboration (AGPC) benchmark, a spatio-temporally aligned benchmark comprising more than 745K raw video frames. Built upon this benchmark, we propose Socialized Co-Perception (SCP), a coarse-to-fine framework that organizes collaboration progressively from aerial global localization to ground target association and identity-aware parsing. Its core module, the Dual-Layer Router (DLR), decouples input-side multi-scale expert selection from output-side task-conditioned modulation, enabling selective cross-view and cross-task interaction while suppressing harmful interference. Extensive experiments demonstrate the effectiveness of SCP. It achieves a 3.73\% coevolutionary gain and a 7.86\% improvement in average downstream performance. These results show that task-conditioned collaboration is more effective than uniform fusion for heterogeneous air-ground perception. The code is available at https://github.com/g1136639260-spec/AGSCP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。