arXiv:2510.19139cs.AIcs.CL2025-10

测试大模型评估临床试验报告合规性的能力,发现其判断常不靠谱。

A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist

  • 用三种提示策略对比通用与医疗专用大模型的推理表现
  • 两模型在临床角色模拟下校准误差超标,严重高估自身准确性
  • 强调需改进模型校准、透明代码和提示工程以提升医疗AI可靠性

尽管大语言模型(LLMs)在医疗领域快速应用,但对其评估临床试验报告是否符合 CONSORT 标准的能力进行稳健且可解释的评测仍是未解难题。特别是,模型推理中的不确定性校准和元认知可靠性在医疗自动化中尚未得到充分理解与探索。本研究采用专家验证的数据集,通过行为与元认知分析方法,系统比较了两种代表性大模型——一个通用模型和一个领域专精模型——在三种提示策略下的表现。我们使用期望校准误差(ECE)和基线归一化的相对校准误差(RCE)来分析认知适应性和校准误差,实现跨模型可靠比较。结果表明,两种模型在临床角色扮演条件下均存在显著校准偏差和过度自信,校准误差持续高于临床相关阈值。这些发现凸显了改进模型校准、增强代码透明度及优化提示工程的必要性,以构建可信赖且可解释的医疗人工智能系统。

原文摘要 · Abstract (English)

Despite the rapid expansion of Large Language Models (LLMs) in healthcare, robust and explainable evaluation of their ability to assess clinical trial reporting according to CONSORT standards remains an open challenge. In particular, uncertainty calibration and metacognitive reliability of LLM reasoning are poorly understood and underexplored in medical automation. This study applies a behavioral and metacognitive analytic approach using an expert-validated dataset, systematically comparing two representative LLMs - one general and one domain-specialized - across three prompt strategies. We analyze both cognitive adaptation and calibration error using metrics: Expected Calibration Error (ECE) and a baseline-normalized Relative Calibration Error (RCE) that enables reliable cross-model comparison. Our results reveal pronounced miscalibration and overconfidence in both models, especially under clinical role-playing conditions, with calibration error persisting above clinically relevant thresholds. These findings underscore the need for improved calibration, transparent code, and strategic prompt engineering to develop reliable and explainable medical AI.

大模型评估医疗AI提示工程校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。