arXiv:2606.08036cs.IRcs.AI2026-06被引 1

测试大模型在地理信息科学中的过度自信,发现其输出常虚假确定。

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

论文配图:GIScholarBench: Benchmarking LLM Overconfidence in GIS Research
图 1 · 摘自论文原文
  • 构建10,865篇论文的基准测试,涵盖三类认知复杂任务
  • 所有模型在错误时仍给出确定答案,文献链接存在不可靠扩增
  • 生成研究方向时覆盖不全、创新点遗漏多,适合科研审慎使用

大型语言模型(LLMs)在学术研究中应用日益广泛,但学术任务要求高事实精度,暴露了其关键弱点:过度自信。此处过度自信定义为行为特征——即使知识不完整或无法验证,仍产生自信、肯定且格式良好的输出,而非仅指置信度与准确率之间的校准偏差。为检验该问题,我们引入GIScholarBench,基于2020至2025年间25本核心地理信息科学期刊发表的10,865篇论文构建。该基准涵盖三类认知复杂度递增的任务:元数据检索、文献关联和研究方向生成。我们在真实用户界面条件下评估Claude Sonnet 4.5、Gemini 3和ChatGPT 5.3。结果显示,所有任务均存在持续过度自信。在元数据检索中,ChatGPT 5.3准确率最高,但所有模型在预测错误时仍给出确定的标题和DOI;在文献关联中,Claude Sonnet 4.5召回参考文献最多,但最高等级检索结果与长列表间存在明显差距,表明参考文献被扩展至可靠检索能力之外;在研究方向生成中,AI生成的方向主题覆盖率较低,新颖性遗漏率较高,语义多样性低于真实未来引用论文。这些发现表明,大模型的过度自信具有任务不变性,但表现形式不同:元数据检索中表现为事实性过度生成,文献关联中体现为不可靠的引文扩展,研究构想中则表现为对输出完整性过度自信。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore expose a key weakness: overconfidence. Here, overconfidence is defined behaviorally as the tendency to produce confident, assertive, and well-formatted outputs even when the underlying knowledge is incomplete or unverifiable, rather than as a calibration gap between stated confidence and accuracy. To examine this issue, we introduce GIScholarBench, a benchmark built from 10,865 papers published in 25 core GIScience journals between 2020 and 2025. The benchmark covers three tasks with increasing cognitive complexity: metadata retrieval, literature linking, and research direction generation. We evaluate Claude Sonnet 4.5, Gemini 3, and ChatGPT 5.3 through their native web interfaces under real-world user-facing conditions. Results show consistent overconfidence across all tasks. In metadata retrieval, ChatGPT 5.3 achieves the highest accuracy, but all models still generate definitive titles and DOIs when predictions are wrong. In literature linking, Claude Sonnet 4.5 recovers the most references, but all models show a clear gap between top-ranked retrieval and longer citation lists, suggesting that references are extended beyond reliable retrieval capacity. In research direction generation, AI-generated directions show lower topic coverage, higher novel miss rates, and lower semantic diversity than real future-citing papers. These findings suggest that LLM overconfidence is task-invariant but takes different forms: factual overgeneration in retrieval, unreliable citation expansion in literature linking, and overconfidence in output completeness during research ideation.

大模型评测过度自信地理信息学术辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。