提出SkyVLaM模型,提升无人机遥感视频中的目标语义理解能力
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

- 通过时序基感知器构建稀疏视觉令牌,融合动态视角信息
- 在101个视频、1.53百万像素级目标上实现更精准的条件分割
- 适合需要高精度无人机视频分析的研究者与应用开发者
近年来,多模态大语言模型(MLLMs)显著提升了遥感多模态理解能力。基于语言的细粒度目标分割在无人机(UAV)视频中至关重要,但因小目标、视觉模糊及动态空中视角仍具挑战。本文提出SkyVLaM,一种面向无人机视频理解的多模态大语言模型。SkyVLaM通过时序基感知器从帧级视频表示中直接构建稀疏令牌,对稀疏基进行正则化以促进互补时序线索,并自适应选择时间连贯的密集片段用于高分辨率检查。稀疏与密集令牌共同输入大语言模型,实现查询条件下的分割。我们进一步构建了SkyVid数据集,包含SkyVid-VGCG(视频基础对话生成)与SkyVid-RVOS(指代视频目标分割),涵盖101个视频、33.6万帧和153万像素级目标实例。实验表明,SkyVLaM在无人机场景下实现了更优的视觉令牌预算分配,显著提升了语言条件视频分割性能。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved remote sensing (RS) multimodal understanding. Language-conditioned segmentation is crucial for fine-grained target understanding in Unmanned Aerial Vehicle (UAV) videos. However, this task remains challenging due to the prevalence of small, visually ambiguous targets and dynamic aerial perspectives. In this paper, we propose SkyVLaM, a multimodal large language model for UAV video understanding. SkyVLaM constructs sparse tokens directly from patch-level video representations through a temporal basis perceiver, regularizes the sparse basis to encourage complementary temporal cues, and adaptively selects a temporally coherent dense segment for high-resolution inspection. The resulting sparse and dense tokens are jointly processed by a large language model for query-conditioned segmentation. We further build SkyVid, consisting of SkyVid-VGCG and SkyVid-RVOS for video grounded conversation generation and referring video object segmentation, respectively. SkyVid contains 101 videos, 33.6K frames, and 1.53M pixel-level object instances. Experiments show that SkyVLaM provides a more effective allocation of the visual token budget and improves language-conditioned video segmentation in UAV scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。