arXiv:2608.02039cs.CV2026-08

为遥感视频设计新数据集与评估基准,提升视觉语言模型理解能力。

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

论文配图:RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?
图 1 · 摘自论文原文
  • 构建10.77万条遥感视频样本,支持连续时空理解评测
  • 提出新框架使模型在小目标追踪上准确率最高达40.63%
  • 适合遥感、多模态、长时序分析方向研究者参考

遥感视频可实现对目标属性、短期行为及场景演化的实时观测,记录了孤立图像无法捕捉的运动、动作、交互与变化。现有模型主要针对单张图像或长时间跨度的离散时序观测,但缺乏统一的连续遥感视频理解评估体系。本文提出RSVideo-10K数据集,包含10,773个实例、147万帧和17.02小时视频,涵盖无人机与卫星平台。其固定评估基准RSVideo-Bench包含2,731个测试实例,评估两个互补维度:L1感知与L2推理,覆盖七个能力组与17项任务。评估表明当前视觉语言模型仍难以恢复微小局部证据、追踪短暂状态及利用场景约束的空间关系。基于此分析,我们进一步提出RSVideo框架,一种基于强化学习的小目标时空聚焦方法,跨帧选择与问题相关的区域并抑制冗余背景标记。该框架在InternVL3.5-14B上实现最大9.01%的绝对提升,在26个开源视觉语言骨干模型中达到最高40.63%的准确率。代码将公开于https://github.com/HongjieZhou0329/RSVideo。

原文摘要 · Abstract (English)

Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene changes that cannot be captured by isolated images. Existing models primarily target single images or discrete temporal observations spanning a long time range. However, a unified evaluation setting for assessing vision-language models on continuous remote-sensing video understanding remains lacking. We introduce RSVideo-10K, a remote-sensing video dataset comprising 10,773 instances, 1.47 million frames, and 17.02 hours of footage, containing both unmanned aerial vehicles and satellite platforms. Its fixed evaluation benchmark, RSVideo-Bench, contains 2,731 test instances and evaluates two complementary aspects of remote-sensing video understanding: L1 Perception and L2 Reasoning, spanning seven capability groups and 17 tasks. Evaluations show that current vision-language models still struggle to recover small local evidence, track short-lived states, and use scene-constrained spatial relations. Based on this analysis, we further propose RSVideo, a reinforcement learning framework for small-target spatiotemporal focusing that selects question-relevant regions across frames and suppresses redundant background tokens. RSVideo achieves a maximum absolute improvement of 9.01% with InternVL3.5-14B and attains the highest accuracy of 40.63% with Qwen3.6-27B across 26 open-source vision-language backbones. Codes will be available at https://github.com/HongjieZhou0329/RSVideo.

遥感视频多模态强化学习小目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。