arXiv:2608.27971cs.CV2026-08

GAAT通过几何感知对齐提升无人机多模态感知的跨模态匹配精度。

GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception

论文配图:GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception
图 1 · 摘自论文原文
  • 先对齐后融合:利用几何先验预测局部对应可靠性,指导稀疏跨模态交互。
  • 在6个下游任务中性能领先,显著优于现有方法。
  • 适用于复杂飞行场景下的多传感器无人机感知,适合系统级诊断与优化。

无人机多模态感知融合可见光(RGB)、红外(IR)、合成孔径雷达(SAR)和深度传感器,在多种环境下实现场景理解。然而,光学特性、分辨率和安装差异常使实际系统仅能实现全局或图像中心对齐。分块后,视差、平台运动和镜头畸变会导致不同模态对应块中心偏移,削弱密集对比学习和跨模态融合所依赖的空间对应关系。本文提出GAAT(Geometry-Aware Alignment Transformer),一种对齐优先的预训练模型,在跨模态交互前估计局部对应可靠性。GAAT引入syncPATC,通过同步视图变换学习块中心一致性,无需对应标注。其输出包括令牌与查询置信度、查询中心及子块偏移等几何先验,识别出残差错位下的可靠局部锚点。基于这些先验,MG-Sparse-MMA在前K_s个可靠区域执行查询驱动的稀疏融合,以几何校准的局部更新替代密集全块交互。RA-QCGCL通过可靠块-块、块-查询、查询-查询的对比分支,使预训练监督与稀疏查询瓶颈对齐。我们构建UAVMeta与StateBench,提供四个源自平台遥测与图像统计的采集状态指标:相机可靠性、观测尺度、视角稳定性与飞行机动复杂度。大量实验在六个下游任务中验证了其一致优异的迁移性能,确立了GAAT作为无人机感知的先进多模态基础模型地位。StateBench进一步实现对真实采集条件的系统性诊断。

原文摘要 · Abstract (English)

Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in optics, resolution, and mounting often limit practical systems to global or image-center alignment. After tokenization, parallax, platform motion, and lens distortion can shift corresponding patch centers across modalities, weakening the spatial correspondence assumed by dense contrastive learning and cross-modal fusion. We propose GAAT (Geometry-Aware Alignment Transformer), an alignment-first pretrained model that estimates local correspondence reliability before cross-modal interaction. GAAT introduces syncPATC, which learns patch-center consistency under synchronized view transformations without correspondence annotations. It emits geometric priors, including token and query confidence, query centers, and sub-token offsets, that identify reliable local anchors across residual misalignment. Guided by these priors, MG-Sparse-MMA performs query-mediated sparse fusion over top-K_s reliable regions, replacing dense all-patch interaction with geometry-calibrated local updates. RA-QCGCL aligns pretraining supervision with this sparse query bottleneck through reliable patch-to-patch, patch-to-query, and query-to-query contrastive branches. We introduce UAVMeta and StateBench, which provide four acquisition-state scores derived from platform telemetry and image statistics: camera reliability, observation scale, viewpoint stability, and flight maneuver complexity. Extensive experiments across six downstream tasks demonstrate consistently superior transfer performance, establishing GAAT as a state-of-the-art multimodal foundation model for UAV perception. StateBench further enables a systematic diagnosis of real-world acquisition conditions.

多模态感知无人机几何对齐稀疏融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。