arXiv:2606.28369cs.IRcs.AI2026-06中稿 · publication in the…

用AI模型融合图文信息,精准搜索环境事件的时空相似报告

Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models

论文配图:Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models
图 1 · 摘自论文原文
  • 结合文本与图像生成更丰富的事件嵌入表示
  • 引入多尺度时空相关性,显著提升推荐排序效果
  • 适合关注环境变化监测与地理信息智能分析的研究者

针对包含空间与时间信息的异常环境事件(如阿拉斯加搁浅鲸鱼)新闻和报告的语义检索与推荐,是地理信息检索(GIR)中的关键任务。本文提出一种新框架,利用大语言模型(LLMs)与视觉-语言模型(VLMs)实现有效相似性搜索与排序。为此,提出两项新策略:(1) CAMERA(上下文感知多模态事件检索算法),融合文本与视觉信息,生成比仅依赖文本更丰富的嵌入;(2) ASTRA(自适应时空重排序算法),通过引入与尺度相关的时空相关性,优化相似性排序。基于本地环境观察网络数据集的实验表明,该VLM增强方法在相似性排序效果上优于单一模态、仅基于LLM的方法。该框架可自动关联相关事件报告,帮助数据管理者与公众深入理解环境变化及其局部影响。结果表明,AI基础模型能通过整合空间、时间、尺度与语义等地理核心概念,推动GIR向多维度智能分析发展。

原文摘要 · Abstract (English)

Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.g., a dead whale washed ashore in Alaska) that contain spatial and temporal information, is a critical task in Geographic Information Retrieval (GIR). This work presents a novel framework that leverages AI foundation models, including Large Language Models (LLMs) and Vision-Language Models (VLMs), to enable effective similarity search and ranking for such event documents. To support this goal, we introduce two new strategies: (1) CAMERA (Context-Aware Multimodal Event Retrieval Algorithm), which fuses textual and visual information to generate richer embeddings than those derived from text alone; and (2) ASTRA (Adaptive Spatial and Temporal Re-ranking Algorithm), which improves similarity ranking by incorporating scale-dependent spatiotemporal relevance alongside semantic similarity. Experimental results, using a dataset from the Local Environmental Observer Network, demonstrate that our VLM-enhanced methods outperform unimodal, LLM-based approaches in similarity ranking effectiveness. By automatically linking relevant event reports, the proposed framework helps both data curators and the general public gain deeper insights into environmental change and its localized impacts. These findings highlight the potential of AI foundation models to advance GIR through multifaceted, intelligent analysis that integrates key geographic concepts: space, time, scale, and semantics.

语义搜索多模态时空分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。