arXiv:2509.18439cs.CLcs.AI2025-09

用AI分析医患对话,自动评估共同决策水平。

Developing an AI framework to automatically detect shared decision-making in patient-doctor conversations

  • 基于对话对齐分数和预训练模型,自动量化医患互动中的共同决策程度。
  • 无风格策略的深度学习模型召回率达0.640,最优模型与决策冲突、患者参与度显著相关。
  • 结果可解释,适合临床研究者评估医疗决策支持工具效果。

共同决策(SDM)是实现以患者为中心护理的关键。目前尚无方法可大规模自动测量SDM。本研究旨在通过语言建模与对话对齐(CA)分数,开发一种自动化测量方法。从一项多中心随机试验中获取157段患者-医生对话视频,共转录为42,559个句子。采用上下文-响应对与负采样训练深度学习(DL)模型,并通过下一句预测(NSP)任务微调BERT模型。每个表现最佳的模型生成四种类型的CA分数。采用随机效应模型(按医师分层),调整年龄、性别、种族和试验组后,评估CA分数与SDM结果(决策冲突量表DCS、观察患者决策参与度12项量表OPTION12)的关系。使用Benjamini-Hochberg法校正多重比较。157名患者中女性占34%,平均年龄70岁(标准差10.8)。医生平均发言字数(1911)高于患者(773)。无风格策略的DL模型召回@1为0.227,而微调后的BERTbase(110M)达到最高召回@1(0.640)。无风格策略的绝对最大值CA(AbsMax,18.36,SE7.74,p=0.025)与最大值CA(Max CA,21.02,SE7.63,p=0.012)与OPTION12相关。微调后的BERTbase(110M)生成的最大值CA与DCS显著相关(-27.61,SE12.63,p=0.037)。模型大小不影响CA分数与SDM的关联性。本研究提出一种可解释、可扩展的自动化方法,用于在真实医患对话中测量SDM,具备大规模评估决策支持策略的潜力。

原文摘要 · Abstract (English)

Shared decision-making (SDM) is necessary to achieve patient-centred care. Currently no methodology exists to automatically measure SDM at scale. This study aimed to develop an automated approach to measure SDM by using language modelling and the conversational alignment (CA) score. A total of 157 video-recorded patient-doctor conversations from a randomized multi-centre trial evaluating SDM decision aids for anticoagulation in atrial fibrillations were transcribed and segmented into 42,559 sentences. Context-response pairs and negative sampling were employed to train deep learning (DL) models and fine-tuned BERT models via the next sentence prediction (NSP) task. Each top-performing model was used to calculate four types of CA scores. A random-effects analysis by clinician, adjusting for age, sex, race, and trial arm, assessed the association between CA scores and SDM outcomes: the Decisional Conflict Scale (DCS) and the Observing Patient Involvement in Decision-Making 12 (OPTION12) scores. p-values were corrected for multiple comparisons with the Benjamini-Hochberg method. Among 157 patients (34% female, mean age 70 SD 10.8), clinicians on average spoke more words than patients (1911 vs 773). The DL model without the stylebook strategy achieved a recall@1 of 0.227, while the fine-tuned BERTbase (110M) achieved the highest recall@1 with 0.640. The AbsMax (18.36 SE7.74 p=0.025) and Max CA (21.02 SE7.63 p=0.012) scores generated with the DL without stylebook were associated with OPTION12. The Max CA score generated with the fine-tuned BERTbase (110M) was associated with the DCS score (-27.61 SE12.63 p=0.037). BERT model sizes did not have an impact the association between CA scores and SDM. This study introduces an automated, scalable methodology to measure SDM in patient-doctor conversations through explainable CA scores, with potential to evaluate SDM strategies at scale.

医疗AI对话分析共同决策自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。