arXiv:2601.06781cs.HCcs.AI2026-01

用手机拍照片自动生成带解说的导览,让旅行更懂你。

AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs

  • 结合照片视觉特征与附近地理信息,智能匹配地标
  • 无需训练,直接在原图上标注并生成图文音频解说
  • 适合想深度了解景点的旅行者和摄影爱好者

我们提出AutoTour,一种通过智能手机自动为用户拍摄的照片生成细粒度地标标注与描述性叙事的系统。其核心思想是将照片中提取的视觉特征与从开放匹配数据库查询的邻近地理空间特征进行融合。与依赖预定义内容或专有数据集的现有导览应用不同,AutoTour利用开放且可扩展的数据源,实现可扩展、上下文感知的基于照片的引导。为此,我们设计了一个无需训练的流程:首先根据用户GPS位置提取并筛选相关地理空间特征;然后通过视觉语言模型(VLM)检测照片中的主要地标,并将其投影到水平空间平面;最后采用几何匹配算法,依据估算的距离与方向将照片特征与对应地理实体对齐。匹配后的特征被直接锚定在原始照片上,并附带大语言模型生成的文本与音频描述,提供类导览体验。实验证明,AutoTour能为知名及冷门地标生成丰富且可解释的标注,开创了一种融合视觉感知与地理理解的交互式、上下文感知探索新范式。

原文摘要 · Abstract (English)

We present AutoTour, a system that enhances user exploration by automatically generating fine-grained landmark annotations and descriptive narratives for photos captured by users. The key idea of AutoTour is to fuse visual features extracted from photos with nearby geospatial features queried from open matching databases. Unlike existing tour applications that rely on pre-defined content or proprietary datasets, AutoTour leverages open and extensible data sources to provide scalable and context-aware photo-based guidance. To achieve this, we design a training-free pipeline that first extracts and filters relevant geospatial features around the user's GPS location. It then detects major landmarks in user photos through VLM-based feature detection and projects them into the horizontal spatial plane. A geometric matching algorithm aligns photo features with corresponding geospatial entities based on their estimated distance and direction. The matched features are subsequently grounded and annotated directly on the original photo, accompanied by large language model-generated textual and audio descriptions to provide an informative, tour-like experience. We demonstrate that AutoTour can deliver rich, interpretable annotations for both iconic and lesser-known landmarks, enabling a new form of interactive, context-aware exploration that bridges visual perception and geospatial understanding.

智能导览多模态地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。