arXiv:2507.20764cs.CV2025-07被引 2

首个面向无人机多模态图像配准的基准数据集,解决真实复杂场景下的配准难题。

ATR-UMMIM: A Benchmark Dataset for UAV-Based Multimodal Image Registration under Complex Imaging Conditions

  • 构建包含可见光、红外等三模态图像的7969组配对数据
  • 覆盖80-300米飞行高度、0-75°视角及全天候光照变化
  • 提供像素级真值与7.8万+目标边界框,支持模型评估与下游任务

多模态融合已成为无人机目标检测的关键技术,各模态提供互补特征以实现鲁棒特征提取。然而,由于不同模态间存在显著的分辨率、视场角和感知特性差异,准确配准是融合前的前提。尽管其重要性突出,目前尚无公开的基准数据集专门针对无人机航拍场景下的多模态图像配准问题,严重制约了先进配准方法在真实环境中的发展与评估。为此,我们提出ATR-UMMIM,首个专为无人机应用设计的多模态图像配准基准数据集。该数据集包含7,969组原始可见光、红外及精确配准的可见光图像,涵盖80至300米飞行高度、0至75°相机角度,以及全天候、全年的光照与天气变化。为保证配准质量,我们设计半自动化标注流程,为每组图像引入可靠的像素级真值。此外,每组数据均标注六项成像条件属性,支持在真实部署环境下评估配准鲁棒性。为进一步支持下游任务,我们在所有配准图像上提供目标级标注,覆盖11类物体,共77,753个可见光与78,409个红外边界框。我们认为ATR-UMMIM将成为推动真实无人机场景中多模态配准、融合与感知发展的基础基准。数据集可从https://github.com/supercpy/ATR-UMMIM下载。

原文摘要 · Abstract (English)

Multimodal fusion has become a key enabler for UAV-based object detection, as each modality provides complementary cues for robust feature extraction. However, due to significant differences in resolution, field of view, and sensing characteristics across modalities, accurate registration is a prerequisite before fusion. Despite its importance, there is currently no publicly available benchmark specifically designed for multimodal registration in UAV-based aerial scenarios, which severely limits the development and evaluation of advanced registration methods under real-world conditions. To bridge this gap, we present ATR-UMMIM, the first benchmark dataset specifically tailored for multimodal image registration in UAV-based applications. This dataset includes 7,969 triplets of raw visible, infrared, and precisely registered visible images captured covers diverse scenarios including flight altitudes from 80m to 300m, camera angles from 0° to 75°, and all-day, all-year temporal variations under rich weather and illumination conditions. To ensure high registration quality, we design a semi-automated annotation pipeline to introduce reliable pixel-level ground truth to each triplet. In addition, each triplet is annotated with six imaging condition attributes, enabling benchmarking of registration robustness under real-world deployment settings. To further support downstream tasks, we provide object-level annotations on all registered images, covering 11 object categories with 77,753 visible and 78,409 infrared bounding boxes. We believe ATR-UMMIM will serve as a foundational benchmark for advancing multimodal registration, fusion, and perception in real-world UAV scenarios. The datatset can be download from https://github.com/supercpy/ATR-UMMIM

无人机多模态图像配准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。