arXiv:2511.06316cs.AI2025-11

用视觉语言模型从新闻中精准定位交通事故,误差小于1公里。

ALIGN: A Vision-Language Framework for High-Accuracy Accident Location Inference through Geo-Spatial Neural Reasoning

  • 融合文本与地图的多模态推理框架,模拟人类空间判断。
  • 将定位误差从10.915公里降至0.593公里,官方数据验证为0.465公里。
  • 无需训练,适合缺乏道路事故数据的发展中国家使用。

在低收入和中等收入国家,公共安全与城市规划常因缺乏准确、位置明确的道路事故数据而受阻。从非结构化文本中提取可靠地理信息,需克服传统文本地理编码工具在多语言环境及模糊地点描述下的局限性。本研究提出ALIGN(基于地理空间神经推理的事故位置推断)框架,旨在通过模拟人类空间推理,从非结构化孟加拉语新闻报道与地图线索中推断精确事故坐标。构建了多阶段自动化处理流程,整合大语言模型提取关键线索,结合视觉语言模型进行地图验证。采用代理式架构,设计包含光学字符识别(OCR)、网格空间扫描及三轮几何投票的迭代推理循环,有效降低视觉幻觉。结果表明,该多模态框架显著优于传统纯文本地理解析基线:在验证集上,平均定位误差由不可用的10.915公里降至0.593公里;与达卡大都会警察官方记录比对,平均误差为0.465公里。研究成果为数据稀缺地区提供了高精度、免训练的自动化事故制图基础,支持基于证据的道路安全政策制定,并推动多模态AI在交通分析中的应用。

原文摘要 · Abstract (English)

In low- and middle-income countries, public safety and urban planning initiatives frequently face a critical shortage of accurate, location-specific road crash data. Extracting reliable geospatial information from unstructured text requires overcoming the limitations of traditional text-based geocoding tools, which often fail in multilingual environments with ambiguous place descriptions. This study introduces ALIGN (Accident Location Inference through Geo-Spatial Neural Reasoning), a vision-language framework designed to emulate human spatial reasoning to infer precise accident coordinates from unstructured Bangla news reports and map-based cues. A multi stage automated pipeline was developed to process diverse textual and visual data, integrating large language models for cue extraction with vision-language models for map verification. Using an agentic architecture, we modelled an iterative reasoning loop that combines Optical Character Recognition (OCR), grid-based spatial scanning, and a 3-run geometric voting method to mathematically isolate and reduce visual hallucinations. The findings highlight that the multimodal ALIGN framework significantly outperforms traditional text-only geoparsing baselines. For example, the proposed system successfully reduced the mean localization error from an unusable 10.915 km to a sub-kilometer precision of 0.593 km on a validation dataset. Furthermore, testing the framework against official Dhaka Metropolitan Police records confirmed its reliability by achieving a mean error of 0.465 km. The results provide a high-accuracy, training-free foundation for automated crash mapping in data-scarce regions, supporting evidence-driven road-safety policymaking and the integration of multimodal AI in transportation analytics.

事故定位多模态地理推理发展中国家

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。