arXiv:2607.05682cs.AI2026-07被引 1

让大模型生成的科学问题可审计,靠的是结构化问题证书。

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

  • 用结构化证书记录问题的机制、假设和可证伪性,确保可审查。
  • 在10个科研主题上优于基线,得分高出0.48分(满分5分)。
  • 适合希望提升科研可信度的AI科学家与审稿人使用。

用于科学发现的大语言模型系统日益参与构想、文献综述、实验设计和报告撰写,但其提出的第一研究问题往往难以审计:看似合理却未暴露机制、可证伪性或假设,不利于科学家核查。我们提出FirstResearch,一种基于第一性原理的研究问题生成框架,核心是结构化的研究问题证书。该证书记录原始定义、假设、机制模型、矛盾点、可证伪假设、最小决定性实验及失败更新规则,使问题在下游执行前即可被检验。在10个大模型代理科研主题上,FirstResearch在深度求索盲评协议下优于受AI共同研究员、Agent Laboratory和AI Scientist-v2启发的提示基线。独立的Gemini-2.5-Flash评委对40个基线包重评后,系统排名保持一致,FirstResearch得分为4.86/5,最强基线为4.38/5,平均得分相关系数达0.865。一次重复消融实验显示,仅保留证书的评分已达4.90/5(深求索)和4.88/5(Gemini),而移除证书后得分低于1/5。结果初步,使用大模型评委而非人类专家,但仍支持一个狭义科学发现主张:显式推导约束是提升大模型生成科学问题可审计性的有效机制。代码、提示、输出和复现脚本见https://github.com/louiswang524/FirstResearch。

原文摘要 · Abstract (English)

LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect. We introduce FirstResearch, a first-principles research-question formation framework for scientific LLM agents whose core artifact is a structured Research Question Certificate. The certificate records primitive definitions, assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule, making the proposed question inspectable before downstream execution. On ten LLM-agent research topics, FirstResearch outperforms controlled prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2 under a primary DeepSeek-blind-judge protocol. A Gemini-2.5-Flash independent-judge rescore of the same 40 baseline packages preserves the system-level ranking, with FirstResearch scoring 4.86/5 versus 4.38/5 for the strongest baseline and Pearson agreement of 0.865 on average score. A one-repeat ablation checkpoint further suggests that the certificate-centered core is the strongest component: certificate-only scoring reaches 4.90/5 under DeepSeek and 4.88/5 under Gemini, while removing certificates drops below 1/5 under both judges. These results are preliminary and use LLM judges rather than human domain experts, but they support a narrow scientific-discovery claim: explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable. Code, prompts, saved outputs, and reproduction scripts are available at https://github.com/louiswang524/FirstResearch.

科学发现可审计性大模型研究设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。