arXiv:2604.26567cs.CV2026-04被引 2

AirZoo为无人机3D视觉提供大规模真实数据集,解决训练数据稀缺难题。

AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision

论文配图:AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision
图 1 · 摘自论文原文
  • 基于全球摄影测量网格生成可定制飞行轨迹与天气的仿真环境
  • 覆盖22国378个区域,包含城市与自然场景,标注像素级深度与6-DoF位姿
  • 提升主流模型在图像检索、跨视角匹配等任务上的性能上限

尽管数据驱动的3D视觉进展迅速,但航拍几何3D视觉仍因大规模高保真训练数据严重不足而面临挑战。现有基准多偏向地面或物体中心视角,无法涵盖无人机感知中的复杂视角变换与多样环境条件。为此,我们提出AirZoo——一个统一的大规模航拍几何3D视觉数据集与评测基准。AirZoo具备三大特性:1)可扩展生成流程:利用免费的世界级摄影测量3D网格,渲染具有自定义无人机飞行轨迹与可配置天气/光照的广阔户外环境;2)全面场景多样性:迄今最广泛的区域类型覆盖(横跨22个国家的378个区域),系统涵盖高度结构化的城市景观与复杂的非结构化自然环境;3)丰富的几何标注:每帧提供同步的像素级度量深度与精确的6-DoF地理参考位姿,对几何感知学习至关重要。通过三个严格评估任务——航拍图像检索、跨视角匹配和多视角3D重建,我们证明AirZoo可作为强大预训练引擎。在公开及新收集的真实世界基准上进行的大量实验表明,基于AirZoo微调能显著提升当前最优模型(如MegaLoc、RoMa、VGGT和Depth Anything 3)的性能,为航拍空间智能建立新的性能上限。

原文摘要 · Abstract (English)

Despite the rapid progress in data-driven 3D vision, aerial geometric 3D vision remains a formidable challenge due to the severe scarcity of large-scale, high-fidelity training data. Existing benchmarks, predominantly biased toward ground-level or object-centric views, do not account for complex viewpoint transformations and diverse environmental conditions in UAV-based sensing. To bridge this critical gap, we propose AirZoo, a unified large-scale dataset and benchmark for grounding aerial geometric 3D vision. AirZoo possesses three appealing properties: 1) Scalable Generation Pipeline: Leveraging freely available, world-scale photogrammetric 3D meshes, it renders vast outdoor environments with customizable UAV flight trajectories and configurable weather/illumination. 2) Comprehensive Scene Diversity: It provides the most extensive coverage of region types to date (spanning 378 regions across 22 countries), systematically encompassing both highly structured urban landscapes and complex unstructured natural environments. 3) Rich Geometric Annotations: Each frame provides synchronized, pixel-level metric depth and precise 6-DoF geo-referenced poses, essential for geometry-aware learning. Through three rigorous evaluation tracks -- aerial image retrieval, cross-view matching, and multi-view 3D reconstruction -- we demonstrate that AirZoo serves as a powerful pre-training engine. Extensive experiments on both public and newly collected real-world benchmarks reveal that fine-tuning on AirZoo yields substantial performance gains for SoTA models (e.g., MegaLoc, RoMa, VGGT, and Depth Anything 3), establishing a new performance upper bound for aerial spatial intelligence.

3D视觉无人机数据集几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。