arXiv:2506.06377cs.CYcs.LG2025-06

用LLM检测空间计量研究的经济合理性,发现其适合初筛但难做深度判断。

Evaluating Large Language Model Capabilities in Assessing Spatial Econometrics Research

  • 构建28篇论文的真假摘要,让LLM评估变量选择与系数合理性
  • 顶级模型如GPT-4o在变量一致性判断上F1达0.87,但对系数合理性表现不一
  • 模型选择和论文特征交互影响判断准确率,适合辅助初审而非替代人工

本文研究大型语言模型(LLMs)评估空间计量实证研究经济合理性和理论一致性的能力。我们从28篇已发表论文(2005–2024)中生成了原创及故意修改的“反事实”摘要,并由多种LLM进行评估。模型提供定性意见和变量选择、系数合理性、发表适宜性的二分类判断。结果表明,尽管顶级模型(如GPT-4o)在变量选择一致性判断上表现优异(整体F1分数达0.87),但在系数合理性与整体发表适宜性等深层判断上表现差异显著。评估准确性受模型类型、论文特征及其交互作用显著影响,尤其在细微判断上。这凸显了LLM在辅助初步筛查中的优势,以及在深度经济推理上的局限性,暗示其在同行评审中可作为辅助工具,但仍需严格的人工监督。

原文摘要 · Abstract (English)

This paper investigates Large Language Models (LLMs) ability to assess the economic soundness and theoretical consistency of empirical findings in spatial econometrics. We created original and deliberately altered "counterfactual" summaries from 28 published papers (2005-2024), which were evaluated by a diverse set of LLMs. The LLMs provided qualitative assessments and structured binary classifications on variable choice, coefficient plausibility, and publication suitability. The results indicate that while LLMs can expertly assess the coherence of variable choices (with top models like GPT-4o achieving an overall F1 score of 0.87), their performance varies significantly when evaluating deeper aspects such as coefficient plausibility and overall publication suitability. The results further revealed that the choice of LLM, the specific characteristics of the paper and the interaction between these two factors significantly influence the accuracy of the assessment, particularly for nuanced judgments. These findings highlight LLMs' current strengths in assisting with initial, more surface-level checks and their limitations in performing comprehensive, deep economic reasoning, suggesting a potential assistive role in peer review that still necessitates robust human oversight.

LLM评估空间计量同行评审模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。