arXiv:2602.19217cs.CV2026-02被引 2

让遥感图像问答更懂常识,生成更丰富的问题。

Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing

  • 引入外部常识知识三元组,结合图像描述生成问题。
  • 在两个新数据集上优于现有方法,问题更丰富多样。
  • 适合遥感、视觉语言理解等需要常识的场景研究。

随着遥感图像档案的快速发展,针对图像提问成为获取特定信息或执行语义图像检索的有效方式。然而,当前自动生成的问题往往过于简单且基于模板,限制了问答或视觉对话系统在真实场景中的应用。为提升问题的丰富性与多样性,融合图像内容与常识知识,本文提出一种知识感知的遥感视觉问答生成模型(KRSVQG)。该模型从外部知识源引入相关知识三元组以扩展问题内容,同时利用图像描述作为中间表示,确保问题与图像对齐。此外,KRSVQG采用视觉-语言预训练与微调策略,增强其在低数据条件下的适应能力。为评估模型,本文构建了两个知识感知的遥感视觉问答生成数据集:NWPU-300 和 TextRS-300。实验结果(包括定量指标与人工评估)表明,KRSVQG 在生成多样化、图像与领域知识双重约束的问题方面显著优于现有方法。这一工作推动了视觉-语言研究中超越像素的理解,促进了具备视觉基础的人类常识型视觉-语言系统的发展。

原文摘要 · Abstract (English)

With the rapid development of remote sensing image archives, asking questions about images has become an effective way of gathering specific information or performing semantic image retrieval. However, current automatically generated questions tend to be simplistic and template-based, which hinders the deployment of question answering or visual dialogue systems for real-world applications. To enrich and diversify the questions with both image content and commonsense knowledge, we propose a Knowledge-aware Remote Sensing Visual Question Generation model (KRSVQG). The proposed model incorporates related knowledge triplets from external knowledge sources to broaden the question content, while employing image captioning as an intermediary representation to ground questions to the corresponding images. Moreover, KRSVQG utilizes a vision-language pre-training and fine-tuning strategy, enabling the model's adaptation to low data regimes. To evaluate the proposed KRSVQG model, we construct two knowledge-aware remote sensing visual question generation datasets: the NWPU-300 dataset and the TextRS-300 dataset. Evaluations, including metrics and human assessment, demonstrate that KRSVQG outperforms existing methods and leads to rich questions, grounded in both image and domain knowledge. As a key practice in vision-language research, knowledge-aware visual question generation advances the understanding of image content beyond pixels, facilitating the development of knowledge-enriched vision-language systems with vision-grounded human commonsense.

遥感视觉问答常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。