arXiv:2606.08251cs.CYcs.AI2026-06被引 3

AI在科学创新中缺乏提出反常识观点的能力,依赖人类引导才能发挥价值。

Contemporary AI lacks the imagination to diverge or negate in science

论文配图:Contemporary AI lacks the imagination to diverge or negate in science
图 1 · 摘自论文原文
  • 对比非推理与推理型AI,前者生成想法趋同,后者探索更广但均不自发提出否定假设
  • 科学家更倾向采纳与自身研究相似、可信度高的想法,社会科学家容忍风险更高
  • 当前AI评估系统与专家判断一致性弱,需人类反馈训练模型以提升判断力

关于AI将加速科学发现的宏大宣称远超实际证据,而科学家参与的大规模实证研究仍稀缺。本研究对生物学、医学、化学及社会科学领域121,640篇近期预印本进行最大规模评估,邀请6,749位代表性科学家对基于其论文生成的25,139组大语言模型(LLM)设想进行新颖性、可行性、真实性概率和采纳意愿评分。结果呈现三类模式:第一,非推理型LLM生成想法高度趋同,形成“蜂群思维”,而推理模型虽探索更广假设空间,但无一自发提出否定性假设;第二,科学家偏好与自身研究相似的想法,重视可信度而非新颖性,社会科学家比生命科学家更愿承担风险,资深社会科学家批判最严,因其在需要语境感知和理论演化的多维度领域中,模型表现最差;第三,现有自动评价系统(包括LLM作为裁判和前沿模型)与专家判断一致性低,检索增强和科学家角色提示仅带来微弱改善。我们用人类评分数据后训练的Qwen3-14B奖励模型,可捕捉专家品味,较现有模型提升最高达27%,接近人类同行评审的一致性水平。对2010至2025年3900万篇论文的分析显示,自ChatGPT发布后,否定性主张显著减少,想法趋于收敛。基于代理的模拟进一步表明,在饱和领域应格外珍视人类独特性。总体而言,当前科学向AI仍处于需要人类引导的合作者角色。

原文摘要 · Abstract (English)

Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-in-the-loop evidence is scarce. Here we mount the largest evaluation to date, inviting authors of 121,640 recent preprints in biology, medicine, chemistry, and social science to judge large language model (LLM)-generated ideas derived from their own papers. 6,749 representative scientists returned 25,139 rating sets on novelty, feasibility, probability of being true, and favorability of adoption. Three patterns emerge. First, non-reasoning LLMs collapse into a narrow "hivemind" of similar ideas while reasoning models explore a wider hypothesis space, but no model spontaneously proposes null hypotheses, a move humans make more freely. Second, scientists reward ideas resembling their own and prize probability over novelty, though social scientists tolerate risk more than life scientists; senior social scientists are the harshest critics, and their skepticism is earned, as LLMs falter most in pluralistic fields demanding context-aware interpretation and evolving theories. Third, automated evaluators, including LLM-as-a-judge and state-of-the-art (SOTA) models, agree weakly with expert judgment. Retrieval augmentation and scientist persona prompting yield marginal gains. A Qwen3-14B reward model we post-trained on human ratings captures nuances of taste, beats SOTA models by up to 27%, and closes the gap to the consistency of human peer reviewers. An analysis of 39 million papers from 2010 to 2025 links survey findings to macro-level patterns: following ChatGPT's release, null claims are sharply suppressed and ideas contract. Agent-based simulations further suggest that saturated fields should especially prize human uniqueness. For all the hype, today's AI for science remains a collaborator whose imagination and judgment benefit from human grounding.

AI科学大模型创新机制人类监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。