arXiv:2608.14792cs.CLcs.AI2026-08

小样本提示不够用,监督模型更准地识别儿科诊疗中的共同决策行为。

Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters

  • 用监督学习+冻结嵌入比零样本大模型更能捕捉临床对话中的共决行为。
  • 监督模型在患者分组评估下达到0.227的宏观卡帕系数,提升显著。
  • 需注意数据泄露风险,避免标签外泄影响评估结果可靠性。

研究旨在评估大语言模型(LLM)在真实儿科手术决策对话中通过零样本提示检测共同决策(SDM)行为的有效性,并检验监督学习是否带来增益。分析了21段家庭与医生的录音对话(19名患儿,7,566个语句片段,约6.1小时),由训练编码员标注12种SDM行为(人-人宏卡帕=0.695)。比较了零样本本地模型(Qwen 2.5 32B)、基于冻结句嵌入的监督分类器及其逻辑堆叠的结果,在患者分组的外层折叠、内层交叉拟合阈值及患者重采样置信区间下进行评估。结果显示,零样本模型宏观卡帕为0.139(95% CI 0.111–0.164),监督分类器达0.227(0.186–0.262),提升0.088(0.051–0.119);两者堆叠后达0.242(0.198–0.284)。发现多个语料特定的数据泄露路径,如将兄弟姐妹录音分开处理、允许外层保留患者的标签进入下游模型的少样本示例。结论表明,仅靠零样本提示不足以可靠测量SDM行为,且患者级分组无法防止标签泄露,性能敏感于数据划分单位和标签注入位置。外部验证前,结果不可泛化至其他人群、模型、提示或编码体系。

原文摘要 · Abstract (English)

Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested evaluation. Methods: We analyzed 21 audio-recorded outpatient surgical decision encounters (19 unique patients; 7,566 utterance segments; ~6.1 hours) between families of children with multiple long-term conditions and their surgical providers. Trained coders labeled segments for 12 SDM behaviors (human-human macro Cohen's kappa = 0.695). We compared a zero-shot local LLM (Qwen 2.5 32B), a supervised classifier over frozen sentence embeddings, and their logistic stack, under patient-grouped outer folds with inner cross-fitted thresholds and patient-resampled confidence intervals. Results: The zero-shot LLM reached macro kappa = 0.139 (95% CI 0.111-0.164). The supervised classifier reached kappa = 0.227 (0.186-0.262), a paired improvement of 0.088 (0.051-0.119). A logistic stack of the two reached kappa = 0.242 (0.198-0.284). We identified multiple corpus-specific leakage paths, including grouping sibling recordings separately and allowing labels from an outer held-out patient to enter few-shot exemplars used while fitting downstream models. Conclusion: Zero-shot prompting alone is not sufficient to measure SDM behavior as reliably as a small supervised model, and patient-level grouping alone does not prevent leakage when labeled prompt exemplars are precomputed outside the outer evaluation loop. Reported performance is sensitive to the unit of data splitting and to where labeled exemplars enter the pipeline. External validation is needed before these findings generalize beyond this population, model, prompt, and codebook.

临床AI共决识别监督学习数据泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。