AI科研需按证据说话,这篇论文给出了判断可信度的框架。
The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims

- 将AI科研拆解为五步:假设生成、推导后果、验证、更新信念、校准声明
- 提出‘证据许可’概念,强调只有证据支持才能发声,否则就是越权
- 适合关注AI科研可信度的研究者和期刊审稿人
AI辅助研究已进入新阶段,核心问题不再是系统能否生成假说、运行实验或撰写论文,而是其科学主张是否与支撑证据相匹配。本文提出一个概念与方法框架,用于界定AI科研中‘证据许可的主张’。基于科学大模型、LLM研究助手、多智能体协作科学家、AI科学家流水线、数学发现代理及自主实验室等典型路径,将AI科研归纳为五个操作环节:假设生成、模型中介的后果推导、外部验证、信念更新与主张校准。核心观点是,校准并非简单谨慎措辞,而是一种管理科学断言权限的机制——证据决定哪些说法可被允许,哪些不可。文章区分了语言、基于后果、干预性与证据许可四类语义;定义了‘主张-证据差距’与‘认知债务’;并提出对异构输出进行最小结构重构,是主张校准的一种向上形式。附带演示案例AISim-Cal,仅为合成动态示例,非实证预测或基准。得出三大原则:无许可则无主张,验证不决定主张层级,自动化放大校准需求。因此,可靠的AI辅助研究应形成闭环:生成假设、推导可检验后果、接受独立裁决、更新信念,并仅输出经证据许可的主张。
原文摘要 · Abstract (English)
AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether their scientific claims are calibrated to the evidence that supports them. This Perspective-style paper develops a conceptual and methodological framework for evidence-licensed claims in AI-assisted research. Motivated by representative routes including specialized scientific foundation models, LLM research assistants, multi-agent co-scientists, AI Scientist pipelines, mathematical discovery agents, and self-driving laboratories, it represents AI-assisted research as five operators: hypothesis generation, model-mediated consequence derivation, external validation, belief update, and claim calibration. The central claim is that calibration is not merely cautious wording but a mechanism for managing scientific assertion rights: evidence licenses some forms of speech and withholds others. The paper distinguishes linguistic, consequence-based, interventional, and evidence-licensed semantics; defines the claim-evidence gap and epistemic debt; and treats minimal structural reconstruction across heterogeneous outputs as an upward form of claim calibration. AISim-Cal is included as an illustrative synthetic dynamics exercise, not as an empirical forecast or benchmark. The resulting principles are: no claim without license, validation does not determine claim level, and automation amplifies the need for calibration. Reliable AI-assisted research is therefore evaluated as a loop that generates hypotheses, derives testable consequences, accepts independent adjudication, updates beliefs, and outputs only evidence-licensed claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。