构建街景多模态推理数据集,实现精准定位与可解释性说明
GeoExplain: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- 基于街景图像的多层级视觉线索进行跨模态推理
- 提出包含40350组全景图-位置-解释的数据集
- 适用于地理定位与可解释AI研究者
多模态推理是理解、整合并推断不同数据模态信息的过程,近年来受到广泛关注。尽管已有多种任务用于评估多模态推理能力,但仍存在局限性。对不同粒度层级的视觉线索(如局部细节与全局上下文)进行层次化推理,虽在人类认知中常见,却缺乏深入探讨。为此,我们提出一个具有挑战性的数据集GeoExplain,用于评估可解释的地理定位任务。给定一张街景全景图,任务是预测其位置并提供详细解释。GeoExplain包含40350个全景图-位置-解释三元组,每条实例包含一组街景全景图、街道级位置以及由专家撰写的解释,说明如何从图像内容中推断出该位置。此外,我们提出一种多模态多层次推理方法SightSense,能够生成预测结果与完整解释。实验分析表明,该方法在GeoExplain上表现优异。
原文摘要 · Abstract (English)
Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention. Although there are various tasks for evaluating multimodal reasoning ability, they still have limitations. Reasoning on hierarchical visual clues at different levels of granularity, i.e., local details and global context, is of little discussion, despite its frequent involvement in human reasoning. To bridge the gap, we introduce a challenging dataset, namely GeoExplain, which evaluates explainable geo-localization. Given a street view image, the task is to predict its location and provide a detailed explanation. GeoExplain consists of 40350 panoramas-location-explanation tuples. Each instance contains a set of street-view panoramas, a location on street level, and human-expert explanations describing how the location can be inferred from the visual content of panoramas. Additionally, we present a multimodal and multilevel reasoning method, namely SightSense which can make predictions and generate a comprehensive explanation. Our analysis and experiments demonstrate its outstanding performance in GeoExplain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。