arXiv:2512.14102cs.CVcs.AI2025-12被引 3

用符号推理提升遥感图文检索的准确与可解释性

Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries

  • 将文本查询转为逻辑表达式,通过神经符号推理匹配图像实体
  • 在复杂查询下性能超越现有模型,鲁棒性提升30%以上
  • 适合需要高可解释性的灾后遥感图像检索场景

随着面向航空与卫星影像的大规模视觉语言模型(RS-LVLMs)发展,遥感图文检索取得显著进展。然而,可解释性差和对复杂空间关系处理能力不足仍是实际应用的关键挑战。为此,本文提出RUNE(Reasoning Using Neurosymbolic Entities),结合大语言模型(LLM)与神经符号人工智能,通过解析文本查询生成一阶逻辑(FOL)表达式,并在检测到的实体上进行显式推理。不同于依赖隐式联合嵌入的RS-LVLMs,RUNE采用条件子集上的逻辑分解策略,实现更快的推理速度。我们仅用基础模型生成FOL表达式,推理由神经符号模块完成。评估中,我们扩展了原有用于目标检测的DOTA数据集,加入更复杂的查询。结果表明,LLM在文本到逻辑转换中表现有效,且RUNE在复杂遥感检索任务中优于主流模型。我们引入两项新指标:查询复杂度鲁棒性(RRQC)与图像不确定性鲁棒性(RRIU)。RUNE在性能、鲁棒性与可解释性方面均表现更优。通过洪水灾后卫星图像检索案例,验证其在真实场景的应用潜力。

原文摘要 · Abstract (English)

Text-to-image retrieval in remote sensing (RS) has advanced rapidly with the rise of large vision-language models (LVLMs) tailored for aerial and satellite imagery, culminating in remote sensing large vision-language models (RS-LVLMS). However, limited explainability and poor handling of complex spatial relations remain key challenges for real-world use. To address these issues, we introduce RUNE (Reasoning Using Neurosymbolic Entities), an approach that combines Large Language Models (LLMs) with neurosymbolic AI to retrieve images by reasoning over the compatibility between detected entities and First-Order Logic (FOL) expressions derived from text queries. Unlike RS-LVLMs that rely on implicit joint embeddings, RUNE performs explicit reasoning, enhancing performance and interpretability. For scalability, we propose a logic decomposition strategy that operates on conditioned subsets of detected entities, guaranteeing shorter execution time compared to neural approaches. Rather than using foundation models for end-to-end retrieval, we leverage them only to generate FOL expressions, delegating reasoning to a neurosymbolic inference module. For evaluation we repurpose the DOTA dataset, originally designed for object detection, by augmenting it with more complex queries than in existing benchmarks. We show the LLM's effectiveness in text-to-logic translation and compare RUNE with state-of-the-art RS-LVLMs, demonstrating superior performance. We introduce two metrics, Retrieval Robustness to Query Complexity (RRQC) and Retrieval Robustness to Image Uncertainty (RRIU), which evaluate performance relative to query complexity and image uncertainty. RUNE outperforms joint-embedding models in complex RS retrieval tasks, offering gains in performance, robustness, and explainability. We show RUNE's potential for real-world RS applications through a use case on post-flood satellite image retrieval.

遥感检索神经符号可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。