临床大模型虽能隐含表达证据强度,但说出来的话却不可靠。
The strength of clinical evidence is recoverable from language model representations but not from their stated grades

- 从模型激活值中可解码出证据等级,但不如直接问来得准。
- 模型声称的证据等级随机性高,平均比真实等级低25-27个百分点。
- 信号来自词汇特征,不跨主题或框架通用,适合研究模型内部表征的人看。
大型语言模型(LLMs)越来越多地用于总结临床证据,其中主张的支持强度至关重要。然而这些模型在表达信心时表现不佳,且其未明言的属性(如真实性)常可从激活值中读取。本研究测试了临床模型是否能区分并表达证据强度(而非仅真值)。我们从六个公开来源整理出45,134条临床主张,并将20,611条统一为三套独立框架下的四类证据等级。测试了22个本地化、开源权重的LLM(参数量0.6-70亿,涵盖通用、医学及推理类模型),包含词汇、真值与跨框架控制。线性估计器在所有模型中均成功恢复等级(中位数AUROC 71.8),但解码能力不随规模提升,推理类模型最弱。模型自述的等级仅达随机水平,比估计值低25-27个百分点。可恢复信号主要源于词汇特征,无法跨主题或框架迁移,但与事实真值分离,仍能识别弱支持主张(AUROC 69.2)。因此,临床大模型虽隐含有序证据强度信号,却未在输出中表达,导致其声明的等级无法反映真实支持程度。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly summarize clinical evidence, where a claim's weight depends on how strongly it is supported. Yet these models convey confidence poorly, and properties they never state, such as truth, are often readable from their activations. Whether a clinical model registers evidence strength, distinct from truth, and states it when asked is untested, and any such signal could be lexical. We compiled 45,134 clinical claims from six public sources, harmonized 20,611 into a four-level evidence grade under three independent frameworks, and tested 22 local, open-weight LLMs from several developers (0.6-70 billion parameters; general, medical, and reasoning), with lexical, truth, and cross-framework controls. A linear estimator recovered the grade in every model (median AUROC 71.8), yet decodability did not rise with scale and was weakest in reasoning models. The grade the models stated fell to chance, 25-27 percentage points below the estimator. The recoverable signal was largely lexical and did not transfer across topics or frameworks, yet it was distinct from factual truth and still flagged weakly supported claims (AUROC 69.2). Clinical LLMs thus carry an ordered evidence-strength signal they do not express, so their stated grades fail to convey a claim's support even when it is recoverable from their representations and text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。