用大模型自动批评科学模型,提升发现准确度。
CriticAL: Critic Automation with Language Models
- 让大模型生成统计量并用假设检验判断模型与数据差异是否显著
- 在合成数据中准确生成正确批判,不产生虚假错误批判
- 批判结果透明可操作,助大模型改进真实世界模型
通过模型理解世界是科学研究的根本目标。尽管基于大语言模型(LLM)的方法在自动化科学发现方面展现潜力,但往往忽视了批判科学模型的重要性。模型批判能深化科学理解并推动更精确模型的发展。自动化模型批判困难,因其传统上依赖专家定义如何比较模型与数据,并评估偏差是否显著——这高度依赖对建模假设和领域的理解。虽然基于LLM的批判方法吸引人,但存在幻觉风险:模型可能虚构出不存在的批判。为此,我们提出CriticAL(Critic Automation with Language Models)。CriticAL利用LLM生成捕捉模型预测与数据之间差异的摘要统计量,并通过假设检验评估其显著性。可将CriticAL视为验证器,通过将模型及其批判嵌入假设检验框架来验证。实验中,我们在关键定量与定性维度评估CriticAL。在合成模型与数据间存在偏差的场景下,CriticAL可靠地生成正确批判,且未出现错误幻觉。人类与LLM评审员一致认为,CriticAL的批判在透明度与可操作性上优于其他方法。最终,我们证明CriticAL的批判使一个LLM科学家能在真实数据集上改进人类设计的模型。
原文摘要 · Abstract (English)
Understanding the world through models is a fundamental goal of scientific research. While large language model (LLM) based approaches show promise in automating scientific discovery, they often overlook the importance of criticizing scientific models. Criticizing models deepens scientific understanding and drives the development of more accurate models. Automating model criticism is difficult because it traditionally requires a human expert to define how to compare a model with data and evaluate if the discrepancies are significant--both rely heavily on understanding the modeling assumptions and domain. Although LLM-based critic approaches are appealing, they introduce new challenges: LLMs might hallucinate the critiques themselves. Motivated by this, we introduce CriticAL (Critic Automation with Language Models). CriticAL uses LLMs to generate summary statistics that capture discrepancies between model predictions and data, and applies hypothesis tests to evaluate their significance. We can view CriticAL as a verifier that validates models and their critiques by embedding them in a hypothesis testing framework. In experiments, we evaluate CriticAL across key quantitative and qualitative dimensions. In settings where we synthesize discrepancies between models and datasets, CriticAL reliably generates correct critiques without hallucinating incorrect ones. We show that both human and LLM judges consistently prefer CriticAL's critiques over alternative approaches in terms of transparency and actionability. Finally, we show that CriticAL's critiques enable an LLM scientist to improve upon human-designed models on real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。