arXiv:2604.15670cs.CV2026-04中稿 · CVPR被引 2

提出面向无人机图像的多模态推理分割模型PixDLM,解决视角复杂、分辨率高等挑战。

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

  • 构建双路径多模态模型,融合像素级视觉与语言信息进行推理
  • 在包含1万张高分辨率图像的DRSeg数据集上实现强基线性能
  • 适用于需要空间、属性、场景多维度推理的遥感图像分析任务

推理分割已从地面场景扩展至遥感影像,但无人机数据面临俯仰视角、超高分辨率及极端尺度变化等挑战。为此,我们正式定义了无人机推理分割任务,并将其语义需求分为空间、属性和场景三个维度。基于此,构建了包含1万张高分辨率航拍图像的DRSeg大规模基准数据集,每张图像均配有链式思维问答标注,覆盖全部三类推理。作为基准配套模型,提出PixDLM——一种简单而高效的像素级多模态语言模型,可作为该任务的统一基线。在DRSeg上的实验验证了其有效性,揭示了无人机推理分割的独特挑战,为后续研究提供了坚实基础。

原文摘要 · Abstract (English)

Reasoning segmentation has recently expanded from ground-level scenes to remote-sensing imagery, yet UAV data poses distinct challenges, including oblique viewpoints, ultra-high resolutions, and extreme scale variations. To address these issues, we formally define the UAV Reasoning Segmentation task and organize its semantic requirements into three dimensions: Spatial, Attribute, and Scene-level reasoning. Based on this formulation, we construct DRSeg, a large-scale benchmark for UAV reasoning segmentation, containing 10k high-resolution aerial images paired with Chain-of-Thought QA supervision across all three reasoning types. As a benchmark companion, we introduce PixDLM, a simple yet effective pixel-level multimodal language model that serves as a unified baseline for this task. Experiments on DRSeg establish strong baseline results and highlight the unique challenges of UAV reasoning segmentation, providing a solid foundation for future research.

无人机分割多模态模型推理分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。