LLM做实验设计效果差,混合方法更靠谱
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
- 用LLM生成实验建议,但对真实反馈不敏感
- 传统方法如高斯过程优化性能明显更好
- 提出新混合方法,结合先验与最近邻采样
大语言模型(LLMs)近期被宣称可作为通用实验设计代理,具备上下文内实验设计能力。本文通过开放和闭源指令微调的LLMs,在基因扰动和分子性质发现任务中评估该假设。结果发现,基于LLM的代理对实验反馈完全不敏感:将真实结果替换为随机打乱标签后,性能无变化。在多个基准测试中,线性带宽和高斯过程优化等经典方法始终优于LLM代理。为此,我们提出一种简单混合方法——LLM引导的最近邻采样(LLMNN),结合LLM先验知识与最近邻采样以指导实验设计。该方法在无需大量上下文适应的情况下,在各领域均达到竞争力或更优表现。结果表明,当前开放和闭源的LLMs无法真正实现上下文内的实验设计,强调需构建分离先验推理与后验更新的混合框架。
原文摘要 · Abstract (English)
Large language models (LLMs) have recently been proposed as general-purpose agents for experimental design, with claims that they can perform in-context experimental design. We evaluate this hypothesis using both open- and closed-source instruction-tuned LLMs applied to genetic perturbation and molecular property discovery tasks. We find that LLM-based agents show no sensitivity to experimental feedback: replacing true outcomes with randomly permuted labels has no impact on performance. Across benchmarks, classical methods such as linear bandits and Gaussian process optimization consistently outperform LLM agents. We further propose a simple hybrid method, LLM-guided Nearest Neighbour (LLMNN) sampling, that combines LLM prior knowledge with nearest-neighbor sampling to guide the design of experiments. LLMNN achieves competitive or superior performance across domains without requiring significant in-context adaptation. These results suggest that current open- and closed-source LLMs do not perform in-context experimental design in practice and highlight the need for hybrid frameworks that decouple prior-based reasoning from batch acquisition with updated posteriors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。