arXiv:2510.13243cs.CV2025-10被引 3

构建了含真实与合成数据的多模态无人机城市场景数据集

FlyAwareV2: A Multimodal Cross-Domain UAV Dataset for Urban Scene Understanding

  • 融合真实与合成图像,包含RGB、深度和语义标签
  • 提供多种天气和光照条件下的数据,支持跨域适应研究
  • 适合做无人机城市场景理解与模型泛化能力评估的研究者

城市环境中无人机应用的计算机视觉算法发展严重依赖大规模、标注精准的数据集。然而,真实无人机数据的采集与标注成本高昂且困难。为解决这一问题,我们提出FlyAwareV2,一个新型多模态数据集,涵盖真实与合成无人机影像,专用于城市场景理解任务。该数据集在近期发布的SynDrone和FlyAware基础上,新增四大关键贡献:1)覆盖不同天气与昼夜条件的多模态数据(RGB、深度图、语义标签);2)利用先进单目深度估计方法为真实样本生成深度图;3)在标准架构上提供RGB与多模态语义分割基准测试;4)开展从合成到真实的域适应研究,评估模型在合成数据上训练后的泛化能力。凭借丰富的标注与环境多样性,FlyAwareV2为基于无人机的三维城市场景理解研究提供了重要资源。

原文摘要 · Abstract (English)

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating real-world UAV data is extremely challenging and costly. To address this limitation, we present FlyAwareV2, a novel multimodal dataset encompassing both real and synthetic UAV imagery tailored for urban scene understanding tasks. Building upon the recently introduced SynDrone and FlyAware datasets, FlyAwareV2 introduces several new key contributions: 1) Multimodal data (RGB, depth, semantic labels) across diverse environmental conditions including varying weather and daytime; 2) Depth maps for real samples computed via state-of-the-art monocular depth estimation; 3) Benchmarks for RGB and multimodal semantic segmentation on standard architectures; 4) Studies on synthetic-to-real domain adaptation to assess the generalization capabilities of models trained on the synthetic data. With its rich set of annotations and environmental diversity, FlyAwareV2 provides a valuable resource for research on UAV-based 3D urban scene understanding.

无人机多模态数据城市感知域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。