LLM在科学决策中看似稳定,实则常偏离真实结果。
When Stability Fails: Hidden Failure Modes Of LLMS in Data-Constrained Scientific Decision-Making
- 分离评估稳定性、正确性等四维度,精准定位问题
- 相同输入下,模型稳定性高但结果常偏离统计真值
- 适合科研自动化流程的可靠性验证与部署参考
大型语言模型(LLMs)正被广泛用于数据受限的科学工作流中作为决策支持工具,其正确性和有效性至关重要。然而,当前评估常强调重复运行的稳定性或可复现性。尽管这些特性理想,但稳定性本身并不能保证与统计真值的一致性。本文提出一种受控的行为评估框架,明确区分四个决策维度:稳定性、正确性、提示敏感性及固定统计输入下的输出有效性。通过从差异表达分析中提取的基因优先排序任务,在严格与宽松显著性阈值、边界排名场景及微小措辞变化等多种提示条件下评估多个LLM。实验表明,即使模型表现出近乎完美的运行间稳定性,仍可能系统性偏离统计真值:在宽松阈值下过度选择、对微小提示措辞变化反应剧烈,或生成输入表中不存在的语法合理基因标识符。这说明稳定性反映的是重复运行的鲁棒性,而非与统计真值的一致性。研究强调,在自动化或半自动化科学工作流中部署LLM时,必须进行显式的真值验证和输出有效性检查。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as decision-support tools in data-constrained scientific workflows, where correctness and validity are critical. However, evaluation practices often emphasize stability or reproducibility across repeated runs. While these properties are desirable, stability alone does not guar- antee agreement with statistical ground truth when such references are available. We introduce a controlled behavioral evaluation framework that explicitly sep- arates four dimensions of LLM decision-making: stability, correctness, prompt sensitivity, and output validity under fixed statistical inputs. We evaluate multi- ple LLMs using a statistical gene prioritization task derived from differential ex- pression analysis across prompt regimes involving strict and relaxed significance thresholds, borderline ranking scenarios, and minor wording variations. Our ex- periments show that LLMs can exhibit near-perfect run-to-run stability while sys- tematically diverging from statistical ground truth, over-selecting under relaxed thresholds, responding sharply to minor prompt wording changes, or producing syntactically plausible gene identifiers absent from the input table. Although sta- bility reflects robustness across repeated runs, it does not guarantee agreement with statistical ground truth in structured scientific decision tasks. These findings highlight the importance of explicit ground-truth validation and output validity checks when deploying LLMs in automated or semi-automated scientific work- flows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。