arXiv:2603.26177cs.LG2026-03被引 1

AI科学家能通过实验反馈迭代改进科研发现,但需模型能力达到一定门槛。

Can AI Scientist Agents Learn from Lab-in-the-Loop Feedback? Evidence from Iterative Perturbation Discovery

  • 让AI基于实验反馈动态调整假设,实现科学发现的持续优化。
  • 反馈使每特征发现量提升53.4%(p=0.003),且效果依赖真实反馈结构。
  • 模型能力不足时无法有效学习,升级后幻觉率从45%降至9%。

近期研究质疑大语言模型(LLMs)能否在科学实验设计中实现真正的上下文学习(ICL),以往研究认为其对实验反馈不敏感。本文在细胞着色高内涵筛选中开展800次独立复现的迭代扰动发现实验,对比使用实验反馈更新假设的LLM代理与仅依赖预训练知识检索的零样本基线。结果显示,接入反馈使每特征发现量平均提升53.4%(p=0.003)。为检验该提升是否源于真实反馈学习而非提示诱导回忆,引入随机反馈对照组(打乱命中/失败标签),此时性能增益消失(+13.0次命中,p=0.003),表明效果依赖反馈信号结构。进一步分析显示,将Claude Sonnet从4.5升级至4.6,基因幻觉率从约33%-45%降至约3%-9%,使原本不显著的ICL效应(+0.8,p=0.32)转为显著提升(+11.0,p=0.003)。结果表明,有效的实验反馈驱动学习仅在模型达到足够能力阈值后才出现。

原文摘要 · Abstract (English)

Recent work has questioned whether large language models (LLMs) can perform genuine in-context learning (ICL) for scientific experimental design, with prior studies suggesting that LLM-based agents exhibit no sensitivity to experimental feedback. We shed new light on this question by carrying out 800 independently replicated experiments on iterative perturbation discovery in Cell Painting high-content screening. We compare an LLM agent that iteratively updates its hypotheses using experimental feedback to a zero-shot baseline that relies solely on pretraining knowledge retrieval. Access to feedback yields a $+53.4\%$ increase in discoveries per feature on average ($p = 0.003$). To test whether this improvement arises from genuine feedback-driven learning rather than prompt-induced recall of pretraining knowledge, we introduce a random feedback control in which hit/miss labels are permuted. Under this control, the performance gain disappears, indicating that the observed improvement depends on the structure of the feedback signal ($+13.0$ hits, $p = 0.003$). We further examine how model capability affects feedback utilization. Upgrading from Claude Sonnet 4.5 to 4.6 reduces gene hallucination rates from ${\sim}33\%$--$45\%$ to ${\sim}3$--$9\%$, converting a non-significant ICL effect ($+0.8$, $p = 0.32$) into a large and highly significant improvement ($+11.0$, $p=0.003$) for the best ICL strategy. These results suggest that effective in-context learning from experimental feedback emerges only once models reach a sufficient capability threshold.

AI科研反馈学习大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。