AI评估结果快速过时,论文提出可信度追踪框架。
The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research
- 审计40项研究,发现模型平均281天即过时
- 75%研究使用已被替代的模型家族,仅3项更新证据
- 用自反案例展示如何让AI辅助研究可追溯
生成式AI评估在发表前可能已过时,但并非所有结论受时间影响相同。本文对2025年7月18日至2026年7月17日间出现的40项实证记录进行审计,涵盖发表路径、执行时间、模型身份、最新命名模型或不可变快照年龄、同家族替代与刷新行为。新命名模型中位年龄为281天(四分位距:75-478;范围:11-939)。期刊论文中位年龄为395天,预印本为56天,实验室报告为49天。35项包含被替代的模型家族,7项提供精确日期标识,3项明确更新模型证据,1项补充后期敏感性测试。所有40项均包含OpenAI系统,属该语料特征而非统计估计。本文区分模型年龄与主张时效性,提出六项报告规范。其次,将自身两天的创作过程作为前沿模型辅助研究的自反案例。GPT-5.6 Sol Pro在ChatGPT中支持候选发现、源材料核对、计算、起草与批判,作者负责核查来源、做出所有实质性决策并承担责任。此为实践验证,非生产力或质量的控制性估算。通过应用自身提出的模型事实与模型时效声明,论文展示了如何在不将模型输出视为独立验证的前提下,使快速AI辅助研究具备可检查性。标题中的‘半衰期’为比喻,未估算普遍衰减速率。
原文摘要 · Abstract (English)
Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes. First, it audits a maximum-variation purposive corpus of 40 empirical records appearing between 18 July 2025 and 17 July 2026. The audit coded publication route, execution timing, model identity, age of the newest named generation or immutable snapshot, same-family supersession and refresh behaviour. At appearance, the newest named model was a median 281 days old (middle 50%: 75-478; range: 11-939). Median age was 395 days for 25 journal articles, 56 days for 14 preprints and 49 days for one laboratory report. Thirty-five records included a superseded family, seven supplied a precise dated identifier, three clearly refreshed model evidence, and one added a late sensitivity test. All 40 included an OpenAI system, a feature of this corpus rather than a prevalence estimate. The paper distinguishes model age from claim currency and proposes six reporting practices. Second, it treats its own two-day production process as a reflexive case of frontier-model-assisted research creation. GPT-5.6 Sol Pro in ChatGPT supported candidate discovery, source reconciliation, calculations, drafting and critique; the author checked sources, made all substantive decisions and accepts responsibility. This is a proof-of-practice, not a controlled estimate of productivity or quality. By applying its own Model Facts and model-currency statement, the paper shows how rapid AI-assisted research can be made inspectable without treating model output as independent validation. The title uses half-lives metaphorically; no universal decay rate is estimated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。