arXiv:2507.06993cs.AIcs.CV2025-07中稿 · The IEEE/CVF Confe…被引 1

让地图能听懂人话,用自然语言问答和相机视角理解地理信息

IMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial Intelligence

  • 将地图转为可查询的网格表示,支持语言模型理解指代词
  • 融合图像与地理位置信息,实现相机画面到地点的精准定位
  • 轻量多智能体设计,响应快且决策过程可解释

地图应用仍以点选操作为主,难以通过自然语言提问或结合摄像头所见与周围地理上下文。本文提出 IMAIA,一个交互式地图人工智能助手,支持对矢量地图和卫星影像进行自然语言交互,并将摄像头输入与地理空间智能结合,帮助用户理解世界。IMAIA 包含两个互补组件:Maps Plus 将地图切片转换为网格对齐表示,使语言模型能解析指代(如“公园右上角的花形建筑”);Places AI Smart Assistant (PAISA) 通过融合图像-地点嵌入与位置、朝向、距离等地理信号,实现场景定位、突出显著属性并生成简洁解释。采用轻量级多智能体架构,保持低延迟并暴露可解释的中间决策。在地图问答与相机到地点映射任务中,IMAIA 在准确性和响应速度上均优于强基线,同时具备面向用户的实用性。通过统一语言、地图与地理线索,IMAIA 推动地图从脚本化工具迈向空间具身的对话式导航。

原文摘要 · Abstract (English)

Map applications are still largely point-and-click, making it difficult to ask map-centric questions or connect what a camera sees to the surrounding geospatial context with view-conditioned inputs. We introduce IMAIA, an interactive Maps AI Assistant that enables natural-language interaction with both vector (street) maps and satellite imagery, and augments camera inputs with geospatial intelligence to help users understand the world. IMAIA comprises two complementary components. Maps Plus treats the map as first-class context by parsing tiled vector/satellite views into a grid-aligned representation that a language model can query to resolve deictic references (e.g., ``the flower-shaped building next to the park in the top-right''). Places AI Smart Assistant (PAISA) performs camera-aware place understanding by fusing image--place embeddings with geospatial signals (location, heading, proximity) to ground a scene, surface salient attributes, and generate concise explanations. A lightweight multi-agent design keeps latency low and exposes interpretable intermediate decisions. Across map-centric QA and camera-to-place grounding tasks, IMAIA improves accuracy and responsiveness over strong baselines while remaining practical for user-facing deployments. By unifying language, maps, and geospatial cues, IMAIA moves beyond scripted tools toward conversational mapping that is both spatially grounded and broadly usable.

地图智能多模态理解交互式系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。