arXiv:2604.18429cs.CVcs.AI2026-04被引 1

对比两种模型,发现融合更紧密的多模态模型更适合遥感变化问答。

Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models

论文配图:Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models
图 1 · 摘自论文原文
  • 用统一低秩适配框架测试Qwen系列模型
  • 原生多模态模型比分步处理模型表现更好
  • 模型越大不等于越好,融合度比规模更重要

变化视觉问答(Change VQA)旨在回答关于双时相遥感图像语义变化的自然语言问题。尽管视觉语言模型(VLMs)已用于时序遥感图像理解,但针对现代多模态模型的Change VQA仍研究不足。本文在统一的低秩适配(LoRA)设置下,使用最新Qwen模型重新评估CDVQA基准。比较了遵循结构化视觉语言流程的Qwen3-VL(多深度视觉条件+全注意力解码器)与原生多模态模型Qwen3.5(单阶段对齐+混合解码器主干)。在官方CDVQA测试集上的实验表明,近期VLMs优于早期专用基线;但性能并不随模型规模单调提升,且原生多模态模型表现优于结构化视觉语言流程。结果表明,在遥感影像的语言驱动语义变化推理中,紧密集成的多模态主干比规模或显式多深度视觉条件更能提升性能。

原文摘要 · Abstract (English)

Change visual question answering (Change VQA) addresses the problem of answering natural-language questions about semantic changes between bi-temporal remote sensing (RS) images. Although vision-language models (VLMs) have recently been studied for temporal RS image understanding, Change VQA remains underexplored in the context of modern multimodal models. In this letter, we revisit the CDVQA benchmark using recent Qwen models under a unified low-rank adaptation (LoRA) setting. We compare Qwen3-VL, which follows a structured vision-language pipeline with multi-depth visual conditioning and a full-attention decoder, with Qwen3.5, a native multimodal model that combines a single-stage alignment with a hybrid decoder backbone. Experimental results on the official CDVQA test splits show that recent VLMs improve over earlier specialized baselines. They further show that performance does not scale monotonically with model size, and that native multimodal models are more effective than structured vision-language pipelines for this task. These findings indicate that tightly integrated multimodal backbones contribute more to performance than scale or explicit multi-depth visual conditioning for language-driven semantic change reasoning in RS imagery.

遥感多模态视觉问答Qwen

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。