arXiv:2606.12071cs.DLcs.AI2026-06被引 4

LLM评估科研问题新颖性时产生幻觉,与专家意见相悖。

On the Limits of LLM-as-Judge for Scientific Novelty Assessment

论文配图:On the Limits of LLM-as-Judge for Scientific Novelty Assessment
图 1 · 摘自论文原文
  • 用作者锚定的科研问题构建基准,对比模型生成的问题
  • LLM认为生成问题更新颖,但专家更认可原问题
  • 模型常忽略问题的窄域性和来源依赖性,需显式测试

大语言模型(LLM)被广泛用于生成和评判科学想法,新颖性评估成为核心挑战。完整方法评估复杂,因此我们聚焦上游任务:科研问题(RQ)的生成。基于arXiv近期论文,我们构建了RQ-Bench基准,从文献背景、研究空白和贡献中重构作者锚定的RQ。这些问题并非唯一有效,而是用于评测新颖性的参考点。我们采用独立LLM评判、对比式LLM评判及人类专家评估三种方式,发现LLM始终高估生成问题的新颖性,且在对比中偏好更强烈;而领域专家则相反,更倾向作者原始问题。进一步分析显示,许多生成问题过于狭窄或依赖特定来源,而这类维度常被忽视,除非专门测试。整体而言,LLM与专家在新颖性判断上的矛盾,严重质疑了使用LLM评估科研问题新颖性的可靠性。

原文摘要 · Abstract (English)

LLMs are increasingly used to generate and judge scientific ideas. This makes novelty evaluation a central problem. Full idea evaluation is difficult because it often requires judging a method, its feasibility, and its empirical promise. We therefore study a cleaner upstream object: the research question (RQ). RQ generation is a prerequisite for scientific ideation, and RQs can be compared against questions pursued in real papers. We introduce RQ-Bench, a benchmark built from recent arXiv papers. For each paper, we reconstruct author-anchored RQs from its cited background, gaps, and contributions. These RQs are not the only valid questions for the same background. They are author-anchored reference points for testing novelty judgments. We evaluate model-generated RQs with standalone LLM judging, comparative LLM judging, and human expert evaluation. LLM judges consistently rate model-generated RQs as highly novel, producing a novelty mirage; in comparative evaluations, this preference becomes even stronger. Domain experts, however, reach the opposite conclusion and prefer the author-anchored reference questions. We further find that many generated RQs are narrow or source-bound, a dimension that LLM judges often miss unless explicitly tested. Overall, the contradictory novelty evaluations between LLM judges and human experts raise a serious concern about the reliability of using LLMs to assess the scientific novelty of research questions.

LLM评估科研新颖性基准测试专家对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。