arXiv:2605.01011cs.CLcs.AI2026-05被引 1

揭示医学大模型在模糊与噪声下的可靠性下降机制

CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine

论文配图:CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine
图 1 · 摘自论文原文
  • 设计CLEAR框架,系统测试答案数量、真值存在性、表述方式对模型的影响
  • 答案选项越多,模型越难选对或拒绝错误答案,尤其当允许‘我不知道’时出错更多
  • 发现模型越大反而越不谦虚,识别正确答案与拒绝错误答案能力差距扩大

医疗大语言模型评估依赖简化、类似考试的基准,难以反映真实临床询问中的模糊性。我们提出CLinical Evaluation of Ambiguity and Reliability(CLEAR)框架,系统考察决策空间呈现方式、模糊性及不确定性对模型推理的影响。CLEAR通过扰动三方面:(1)合理答案选项数量,(2)是否存在真值或弃权选项,(3)答案选项的语义表述方式。在三个基准上评估17个LLMs发现三大局限:第一,增加合理答案数量会削弱模型识别正确答案及拒绝错误答案的能力;第二,当弃权表述从‘以上皆非’转为‘我不知道’(IDK)时,模型更不愿放弃,且含IDK选项会显著增加错误选择;第三,我们将正确识别与拒绝错误之间的性能差距定义为‘谦逊缺口’,该缺口随模型规模增大而加剧。研究揭示标准医疗基准的缺陷,强调仅靠模型规模无法解决可靠性问题。

原文摘要 · Abstract (English)

Medical large language model (LLM) evaluations rely on simplified, exam-style benchmarks that rarely reflect the ambiguity of real-world medical inquiries. We introduce the CLinical Evaluation of Ambiguity and Reliability (CLEAR) framework, which assesses how decision-space presentation, ambiguity, and uncertainty affect LLMs' reasoning on medical benchmarks. CLEAR systematically perturbs (1) the number of plausible answer options, (2) the presence of a ground truth or abstention option, and (3) the semantic framing of answer options. Applying CLEAR on three benchmarks evaluated across 17 LLMs reveals three notable limitations of existing evaluation methods. First, increasing the number of plausible answers degrades a model's ability to identify the correct answer and abstain against incorrect ones. Second, this lack of caution intensifies as the framing of abstention shifts from assertive rejection like "None of the Above" to uncertainty admission like "I don't know" (IDK). Notably, just including IDK in the answer space increases incorrect answer selections. Lastly, we formalize the performance gap between identifying the correct answer and abstaining from incorrect ones as the humility deficit, which worsens with model scale. Our findings reveal limitations in standard medical benchmarks and underscore that scaling alone does not resolve LLM reliability issues.

医学大模型可靠性评估模糊性谦逊缺口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。