首个评估大模型县区级本地知识的基准,揭示其在地方认知上的严重短板。
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
- 构建覆盖美国49个州526个县的1.48万条问答数据集
- 顶尖模型在叙事类问题上准确率仅56.8%,数值推理低于15.5%
- 模型越大或联网未必更好,体现本地知识理解复杂性
大型语言模型在宏观地理任务(如全球事实召回、事件总结、区域推理)上已广泛评估,但对超本地知识的理解仍不清晰。随着公民平台与社区新闻等应用需求增长,具备邻里动态、文化叙事和地方治理推理能力的AI系统愈发重要。现有基准难以捕捉此类复杂性,多依赖粗粒度数据或孤立信息。本文提出LocalBench,首个系统评估美国县级本地知识的基准。基于本地性概念框架,包含14,782个经验证的问答对,覆盖49个州526个县,融合人口普查数据、本地Reddit讨论与区域新闻。涵盖物理、认知与关系维度的本地性。评估13个先进LLM在闭卷与网络增强两种设置下的表现。结果揭示关键局限:最优模型在叙事类问题上准确率仅为56.8%,数值推理低于15.5%。模型规模扩大或引入网络检索并非必然提升性能——例如,搜索使Gemini准确率提升13.6%,却导致GPT系列下降11.4%。研究凸显亟需能支持公平、地域感知的AI系统,以应对不同地理与文化背景下的精细社区现实。
原文摘要 · Abstract (English)
Large language models (LLMs) have been widely evaluated on macro-scale geographic tasks, such as global factual recall, event summarization, and regional reasoning. Yet, their ability to handle hyper-local knowledge remains poorly understood. This gap is increasingly consequential as real-world applications, from civic platforms to community journalism, demand AI systems that can reason about neighborhood-specific dynamics, cultural narratives, and local governance. Existing benchmarks fall short in capturing this complexity, often relying on coarse-grained data or isolated references. We present LocalBench, the first benchmark designed to systematically evaluate LLMs on county-level local knowledge across the United States. Grounded in the Localness Conceptual Framework, LocalBench includes 14,782 validated question-answer pairs across 526 U.S. counties in 49 states, integrating diverse sources such as Census statistics, local subreddit discourse, and regional news. It spans physical, cognitive, and relational dimensions of locality. Using LocalBench, we evaluate 13 state-of-the-art LLMs under both closed-book and web-augmented settings. Our findings reveal critical limitations: even the best-performing models reach only 56.8% accuracy on narrative-style questions and perform below 15.5% on numerical reasoning. Moreover, larger model size and web augmentation do not guarantee better performance, for example, search improves Gemini's accuracy by +13.6%, but reduces GPT-series performance by -11.4%. These results underscore the urgent need for language models that can support equitable, place-aware AI systems: capable of engaging with the diverse, fine-grained realities of local communities across geographic and cultural contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。