用结构化知识增强无人机视觉问答,提升小物体与关系推理准确率
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning

- 将图像转为包含物体、数量、位置和关系的结构化知识图谱
- 在AUG数据集上超越6个基线模型,尤其在密集场景和关系推理中提升显著
- 适合需要高精度视觉理解的无人机应用与实地部署系统
尽管多模态大语言模型取得进展,但航空场景中的可靠视觉问答仍具挑战。任务关键信息常由小物体、具体数量、粗略位置及物间关系承载,而传统密集视觉标记表示难以匹配这些结构化语义。为此,我们提出AeroRAG,一种基于场景图引导的多模态检索增强生成框架,用于视觉问答。该框架首先将输入图像转化为结构化视觉知识(包括物体类别、数量、空间位置和语义关系),再检索与查询相关的语义片段,构建紧凑提示供给文本型大模型。不依赖对密集视觉标记的直接推理,本方法在感知与语言推理间引入更明确的中间接口。在AUG航空数据集和通用域VG-150基准上,相比六种强基线模型均实现一致提升,尤其在密集航空场景和关系敏感推理中增益最大。进一步在VQAv2上验证,该接口仍兼容标准视觉推理设置。结果表明,结构化检索是面向部署和具身视觉推理系统的可行设计方向。
原文摘要 · Abstract (English)
Despite recent progress in multimodal large language models (MLLMs), reliable visual question answering in aerial scenes remains challenging. In such scenes, task-critical evidence is often carried by small objects, explicit quantities, coarse locations, and inter-object relations, whereas conventional dense visual-token representations are not well aligned with these structured semantics. To address this interface mismatch, we propose AeroRAG, a scene-graph-guided multimodal retrieval-augmented generation framework for visual question answering. The framework first converts an input image into structured visual knowledge, including object categories, quantities, spatial locations, and semantic relations, and then retrieves query-relevant semantic chunks to construct compact prompts for a text-based large language model. Rather than relying on direct reasoning over dense visual tokens, our method introduces a more explicit intermediate interface between perception and language reasoning. Experiments on the AUG aerial dataset and the general-domain VG-150 benchmark show consistent improvements over six strong MLLM baselines, with the largest gains observed in dense aerial scenes and relation-sensitive reasoning. We further evaluate the framework on VQAv2 to verify that the proposed interface remains compatible with standard visual reasoning settings. These results suggest that structured retrieval is a practical design direction for deployment-oriented and grounded visual reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。