arXiv:2511.20460cs.CV2025-11被引 7

不训练即可高效处理超高清遥感图像问答,精准定位关键区域

Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search

  • 通过自适应分枝缩放搜索定位问题相关图像区域
  • 集成后在两个数据集上分别提升114.8%和26.3%准确率
  • 无需训练、即插即用,适合遥感与高分辨率视觉任务

随着卫星星座、传感器技术和成像流程的进步,超高清(Ultra-HR)遥感影像日益普及。然而,现有遥感基础模型难以应对此类输入:全图编码会耗尽标记和内存预算,而缩放预处理则丢失细粒度且影响答案的关键信息。因此,在预测前引导模型关注关键区域变得至关重要。为此,我们提出ZoomSearch——一种无需训练、即插即用的流水线,将‘看哪里’与‘如何回答’解耦,用于超高清遥感视觉问答(RS-VQA)。ZoomSearch结合自适应多分支缩放搜索,对图像块进行层级搜索以定位查询相关区域;以及布局感知块重排,将选定块重组为紧凑且保持布局忠实的画布。我们在Ultra-HR RS-VQA基准MME-RealWorld-RS和LRS-VQA上进行了全面实验,对比了(i)强通用基础模型、(ii)遥感基础模型、(iii)Ultra-HR RS-VQA方法、(iv)即插即用搜索型VQA方法。当与LLaVA-ov集成时,ZoomSearch在多种任务中达到最先进性能,于LRS-VQA上比基线提升26.3%,于MME-RealWorld-RS上提升114.8%。同时,其推理效率更高,比先前搜索方法提速20%~44%。

原文摘要 · Abstract (English)

With advances in satellite constellations, sensor technologies, and imaging pipelines, ultra-high-resolution (Ultra-HR) remote sensing imagery is becoming increasingly widespread. However, current remote sensing foundation models are ill-suited to such inputs: full-image encoding exhausts token and memory budgets, while resize-based preprocessing loses fine-grained and answer-critical details. In this context, guiding the model look where it matters before prediction becomes crucial. Therefore, we present ZoomSearch, a training-free, plug-and-play pipeline that decouples 'where to look' from 'how to answer' for Ultra-HR Remote Sensing Visual Question Answering (RS-VQA). ZoomSearch combines Adaptive Multi-Branch Zoom Search, which performs a hierarchical search over image patches to localize query-relevant regions, with Layout-Aware Patch Reassembly, which reorganizes the selected patches into a compact, layout-faithful canvas. We conduct comprehensive experiments on Ultra-HR RS-VQA benchmarks MME-RealWorld-RS and LRS-VQA, comparing against (i) strong general foundation models, (ii) remote sensing foundation models, (iii) Ultra-HR RS-VQA methods, and (iv) plug-and-play search-based VQA methods. When integrated with LLaVA-ov, ZoomSearch attains state-of-the-art accuracy across diverse tasks, improving the LLaVA-ov baseline by 26.3% on LRS-VQA and 114.8% on MME-RealWorld-RS. Meanwhile, it achieves much higher inference efficiency, outperforming prior search-based methods by 20%~44% in speed.

遥感图像视觉问答超高清高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。