提出首个高分辨率SAR遥感问答数据集,实现雷达与光学图像融合问答。
SAR Strikes Back: A New Hope for RSVQA
- 分两阶段处理:先将SAR图像转为文本再由语言模型作答
- 两阶段模型比端到端提升近10%准确率,融合后整体准确率达75.49%
- 特别擅长水体等特定地物问答,适合遥感与多模态研究者
遥感视觉问答(RSVQA)旨在从卫星图像中提取信息以回答自然语言问题,辅助图像理解。尽管已有多种针对多光谱、多分辨率光学图像的方法,但对高分辨率合成孔径雷达(SAR)图像的研究才刚起步。SAR可在全天气条件下运行并捕捉电磁特征,具有重要潜力,但目前尚无研究比较SAR与光学图像在RSVQA中的表现,也缺乏有效的融合策略。本文构建了首个支持基于SAR的RSVQA数据集,并探索两种处理流程:一是端到端模型,二是两阶段框架——先从SAR图像中提取信息并转为文本,再交由语言模型生成答案。结果表明,两阶段模型表现更优,准确率较端到端方法提升近10%。此外,评估了多种融合策略,决策级融合效果最佳,在所提数据集上达到F1-micro 75.00%、F1-average 81.21%、总体准确率75.49%。SAR尤其在涉及特定地物类型(如水体)的问题上表现突出,验证了其作为光学图像互补模态的价值。
原文摘要 · Abstract (English)
Remote Sensing Visual Question Answering (RSVQA) is a task that extracts information from satellite images to answer questions in natural language, aiding image interpretation. While several methods exist for optical images with varying spectral bands and resolutions, only recently have high-resolution Synthetic Aperture Radar (SAR) images been explored. SAR's ability to operate in all weather conditions and capture electromagnetic features makes it a promising modality, yet no study has compared SAR and optical imagery in RSVQA or proposed effective fusion strategies. This work investigates how to integrate SAR data into RSVQA and how to best combine it with optical images. We present a dataset that enables SAR-based RSVQA and explore two pipelines for the task. The first is an end-to-end model, while the second is a two-stage framework: SAR information is first extracted and translated into text, which is then processed by a language model to produce the final answer. Our results show that the two-stage model performs better, improving accuracy by nearly 10% over the end-to-end approach. We also evaluate fusion strategies for combining SAR and optical data. A decision-level fusion yields the best results, with an F1-micro score of 75.00%, F1-average of 81.21%, and overall accuracy of 75.49% on the proposed dataset. SAR proves especially beneficial for questions related to specific land cover types, such as water areas, demonstrating its value as a complementary modality to optical imagery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。