测试GPT-4o和Gemini在地理信息任务中的表现,发现准确性不一且存在系统偏差。
The World As Large Language Models See It: Exploring the reliability of LLMs in representing geographical features
- 对比GPT-4o与Gemini 2.0 Flash在地理编码、高程估计和反向地理编码中的表现
- 两者均低估奥地利高程,且对联邦州识别存在误判,Gemini略优但仍有明显错误
- 提示需结合地理数据微调,才能提升在地理信息科学中的可靠性
随着大语言模型(LLMs)不断发展,其提供事实信息的可信度日益受到关注,尤其是在地理世界表征方面。本研究评估了GPT-4o与Gemini 2.0 Flash在三个关键地理空间任务中的表现:地理编码、高程估计与反向地理编码。在地理编码任务中,两模型在奥地利因斯布鲁克圣安娜柱坐标估计上均存在系统性与随机误差,GPT-4o偏差更大,而Gemini 2.0 Flash精度更高但有显著系统偏移。在高程估计中,两模型普遍低估奥地利各地高程,尽管整体地形趋势可捕捉,其中Gemini 2.0 Flash在东部区域表现更优。反向地理编码任务(从坐标识别奥地利联邦州)显示,Gemini 2.0 Flash在整体准确率与F1分数上优于GPT-4o,区域一致性更好。然而,两者均未能准确重建奥地利联邦州划分,存在持续误判。研究结论指出,尽管LLMs能近似地理信息,其准确性和可靠性仍不一致,亟需通过地理信息微调以增强其在地理信息科学与地学信息学中的应用价值。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to evolve, questions about their trustworthiness in delivering factual information have become increasingly important. This concern also applies to their ability to accurately represent the geographic world. With recent advancements in this field, it is relevant to consider whether and to what extent LLMs' representations of the geographical world can be trusted. This study evaluates the performance of GPT-4o and Gemini 2.0 Flash in three key geospatial tasks: geocoding, elevation estimation, and reverse geocoding. In the geocoding task, both models exhibited systematic and random errors in estimating the coordinates of St. Anne's Column in Innsbruck, Austria, with GPT-4o showing greater deviations and Gemini 2.0 Flash demonstrating more precision but a significant systematic offset. For elevation estimation, both models tended to underestimate elevations across Austria, though they captured overall topographical trends, and Gemini 2.0 Flash performed better in eastern regions. The reverse geocoding task, which involved identifying Austrian federal states from coordinates, revealed that Gemini 2.0 Flash outperformed GPT-4o in overall accuracy and F1-scores, demonstrating better consistency across regions. Despite these findings, neither model achieved an accurate reconstruction of Austria's federal states, highlighting persistent misclassifications. The study concludes that while LLMs can approximate geographic information, their accuracy and reliability are inconsistent, underscoring the need for fine-tuning with geographical information to enhance their utility in GIScience and Geoinformatics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。