让记者用自然语言搜地理数据,一键验证位置真伪。
SPOT: Bridging Natural Language and Geospatial Search for Investigative Journalists
- 用大模型把口语描述转为精准地理查询
- 在真实新闻场景下准确率超90%且抗干扰
- 专为记者设计,免代码、抗错误输入
OpenStreetMap(OSM)是调查记者进行地理定位验证的重要资源。然而,现有工具如Overpass Turbo需掌握复杂的查询语言,对非技术用户构成障碍。本文提出SPOT,一个开源的自然语言接口,通过直观的场景描述使OSM丰富的标签式地理数据更易获取。SPOT利用微调的大语言模型(LLMs)将用户输入解析为地理对象配置的结构化表示,并在交互式地图界面中展示结果。尽管更通用的地理搜索任务也有可能实现,但SPOT专为调查新闻设计,解决模型幻觉、OSM标签不一致及用户输入噪声等现实挑战。它结合创新的合成数据生成流程与语义聚合系统,实现鲁棒且精确的查询生成。据我们所知,SPOT是首个在此精度水平上实现可靠自然语言访问OSM数据的系统。通过降低地理定位验证的技术门槛,SPOT为事实核查和反虚假信息工作提供了实用工具。
原文摘要 · Abstract (English)
OpenStreetMap (OSM) is a vital resource for investigative journalists doing geolocation verification. However, existing tools to query OSM data such as Overpass Turbo require familiarity with complex query languages, creating barriers for non-technical users. We present SPOT, an open source natural language interface that makes OSM's rich, tag-based geographic data more accessible through intuitive scene descriptions. SPOT interprets user inputs as structured representations of geospatial object configurations using fine-tuned Large Language Models (LLMs), with results being displayed in an interactive map interface. While more general geospatial search tasks are conceivable, SPOT is specifically designed for use in investigative journalism, addressing real-world challenges such as hallucinations in model output, inconsistencies in OSM tagging, and the noisy nature of user input. It combines a novel synthetic data pipeline with a semantic bundling system to enable robust, accurate query generation. To our knowledge, SPOT is the first system to achieve reliable natural language access to OSM data at this level of accuracy. By lowering the technical barrier to geolocation verification, SPOT contributes a practical tool to the broader efforts to support fact-checking and combat disinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。