用精选例子提升大模型临床推理可靠性,实验证明选对例子比堆例子更重要。
Reliability of Large Language Model Generated Clinical Reasoning in Assisted Reproductive Technology: Blinded Comparative Evaluation Study
- 通过精选高质量、多样化的示例,显著提升大模型生成临床推理的准确性。
- 随机示例效果与无示例基本持平,说明低质示例无效。
- 提出双原则框架,适合医疗AI可信数据生成与评估场景。
生成高质量临床思维链(CoTs)对可解释医疗AI至关重要,但受限于数据稀缺。尽管大语言模型(LLMs)能合成医学数据,其临床可靠性尚未验证。本研究评估了LLM生成的CoTs可靠性,并探究提示策略对质量的影响。在一项盲法对比研究中,辅助生殖技术(ART)资深临床医生评估了三种策略生成的CoTs:零样本、随机少量样本(使用浅层示例)、选择性少量样本(使用多样化、高质量示例)。专家评分与先进AI模型GPT-4o的评估结果对比显示,选择性少量样本策略在所有人类评价指标上均显著优于其他策略(p < .001)。关键发现是,随机少量样本策略相比零样本基线无显著改进,表明低质量示例与无示例效果相当。选择性策略的成功归因于两个原则:‘金标准深度’(推理质量)和‘代表性多样性’(泛化能力)。值得注意的是,AI评估器无法识别这些关键差异。合成思维链的临床可靠性取决于提示的策略性筛选,而非示例数量。本文提出‘双原则’框架作为规模化生成可信数据的基础方法。该研究解决了数据瓶颈问题,确认了人类专家在高风险临床AI评估中的不可替代作用。
原文摘要 · Abstract (English)
Creating high-quality clinical Chains-of-Thought (CoTs) is crucial for explainable medical Artificial Intelligence (AI) while constrained by data scarcity. Although Large Language Models (LLMs) can synthesize medical data, their clinical reliability remains unverified. This study evaluates the reliability of LLM-generated CoTs and investigates prompting strategies to enhance their quality. In a blinded comparative study, senior clinicians in Assisted Reproductive Technology (ART) evaluated CoTs generated via three distinct strategies: Zero-shot, Random Few-shot (using shallow examples), and Selective Few-shot (using diverse, high-quality examples). These expert ratings were compared against evaluations from a state-of-the-art AI model (GPT-4o). The Selective Few-shot strategy significantly outperformed other strategies across all human evaluation metrics (p < .001). Critically, the Random Few-shot strategy offered no significant improvement over the Zero-shot baseline, demonstrating that low-quality examples are as ineffective as no examples. The success of the Selective strategy is attributed to two principles: "Gold-Standard Depth" (reasoning quality) and "Representative Diversity" (generalization). Notably, the AI evaluator failed to discern these critical performance differences. The clinical reliability of synthetic CoTs is dictated by strategic prompt curation, not the mere presence of examples. We propose a "Dual Principles" framework as a foundational methodology to generate trustworthy data at scale. This work offers a validated solution to the data bottleneck and confirms the indispensable role of human expertise in evaluating high-stakes clinical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。