arXiv:2604.05377cs.CV2026-04被引 1

让视觉语言模型学会从高空看世界,统一解决无人机场景的推理与生成问题。

Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation

论文配图:Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
图 1 · 摘自论文原文
  • 构建了23.6万张带描述的无人机图像数据集,支持多模态联合建模。
  • 新模型使问答准确率提升至0.973,生成质量显著改善,跨模态一致性增强。
  • 揭示了生成与推理的双向协同机制,适合研究空中智能与多模态生成的学者。

视觉语言模型在地面视角理解上已取得显著进展,但在高空无人机场景中仍表现脆弱:物体微小、密集重叠、纹理重复且俯视方向模糊。为此,我们提出UAVReason,一个大规模、面向无人机的原生数据集与评测套件,用于研究该俯视域下的统一推理与生成任务。UAVReason整合了RGB图像、深度图、语义分割掩码、描述文本及问答对,形成一致的空域数据体系。数据集包含23.6万张带描述的帧、27.3万组VQA问答对(含6.82万组双帧时序问题)以及188.8万组跨模态生成样本,覆盖RGB、深度与分割模态。我们进一步提出UAVReason-Bagel作为统一理解和生成基线,联合优化语言推理与密集视觉生成目标。实验表明,通用视觉语言模型和现成生成器在无人机场景中存在严重定位偏差,而UAVReason-Bagel相比预训练模型大幅改进:单帧问答F1从0.394提升至0.711,双帧问答F1从0.427升至0.822,朝向感知问答F1从0.798增至0.973;生成方面,分割mIoU达0.143,深度-分割-文本条件下的RGB合成KID由0.078降至0.048。更重要的是,消融实验揭示生成与推理之间存在双向协同:密集生成提升时序语义一致性,语言推理则正则化稀疏条件图像生成。结果表明,统一推理与生成可为物理合理的空中智能提供有效的几何结构先验。所有数据、代码与评估工具将公开发布。

原文摘要 · Abstract (English)

Vision-Language Models have achieved strong progress in ground-view visual understanding, yet they remain brittle in high-altitude Unmanned Aerial Vehicle scenes, where objects are tiny and densely packed, textures are repetitive, and top-down orientations are ambiguous. We introduce UAVReason, a large-scale UAV-native dataset and evaluation suite for studying unified aerial reasoning and generation under this nadir-view domain shift. UAVReason aligns RGB imagery, depth maps, semantic segmentation masks, captions, and question-answer pairs within a consistent aerial domain. It contains 23.6K captioned frames, 273K VQA pairs including 68.2K two-frame temporal questions, and 188.8K cross-modal generation samples across RGB, depth, and segmentation modalities. We further adapt UAVReason-Bagel as a unified understanding-and-generation baseline that jointly optimizes language reasoning and dense visual generation objectives. Experiments show that general-purpose VLMs and off-the-shelf unified generators struggle with UAV-native grounding, while UAVReason-Bagel substantially improves over its pretrained counterpart, increasing VQA-1F F1 from 0.394 to 0.711, VQA-2F F1 from 0.427 to 0.822, and heading-aware VQA F1 from 0.798 to 0.973. For generation, it improves segmentation mIoU to 0.143 and reduces KID from 0.078 to 0.048 for depth-segmentation-text-conditioned RGB synthesis. More importantly, our ablations reveal a bidirectional synergy between synthesis and reasoning. Dense generation objectives improve temporal semantic consistency, while language-level reasoning regularizes sparse-condition image synthesis. These results suggest that unified reasoning and generation provide effective geometry-aware structural priors for physically grounded aerial intelligence. All data, code, and evaluation tools will be released.

无人机多模态生成推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。