LLM使用与科研产出正相关,即使排除时间选择偏差仍成立。
A robust association between LLM use and scientific productivity: Assessing stopping-time selection

- 通过多种设计规避时间选择偏差,验证关联稳健性
- 实证发现使用LLM后科研产出显著提升,且高于随机基准
- 适合关注AI对学术影响的研究者和政策制定者
RBB认为,将LLM采用时间定为作者摘要首次被标记的月份,会引发停止时间选择偏差,可能导致在无因果效应时也出现正向事件研究路径。尽管该机制在数学上可行,但不构成零效应的证据。我们将RBB自身的随机安慰剂重新校准至检测器的实际标记率,发现观测到的关联仍远高于该基准,说明该人为效应过小,无法解释生产力变化。我们进一步采用多种互补设计重新估计LLM使用与生产力的关系:跨年前后对比、保守的双重差分控制组、不定义采用日期的强度模型,以及固定标记率的排序测量法。所有设计下均保持正向关联,而同期前ChatGPT的安慰剂数据则显示无效应。RBB识别的偏差确实存在但有限,不足以解释所报告的模式。
原文摘要 · Abstract (English)
Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an adoption date, and a rank-based measurement holding the flag rate fixed. A positive productivity association persists across all of these estimates, while the same tests run on pre-ChatGPT placebo data return null effects. The artifact RBB identify is real but bounded, and it does not account for the pattern we report.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。