arXiv:2608.31007cs.HCcs.CL2026-08中稿 · ACII 2026 as an or…

用自动语言分析辅助医生判断患者诊疗体验,效果优于单独使用任一方法。

Augmenting Interviewer Judgments of Patient Experience with Automatic Language Analysis

  • 融合医生评分与自动语言模型预测,提升对患者体验的评估准确率。
  • 组合模型相关系数达0.403,高于医生单独判断(0.365)和纯自动模型(最高0.286)。
  • 适合临床反馈、医患关系研究及心理治疗质量评估场景。

理解精神科患者对临床对话的主观体验对于反馈和联盟相关过程监测至关重要。尽管访谈者会形成会后对患者体验的判断,但这些判断并不总与患者的自述一致。已有研究提出从对话中自动预测感知互动质量的方法,但尚不清楚这些方法能否补充而非复制人类判断。为填补这一空白,我们评估了一种临床支持框架:将会后访谈者评分与基于语言的自动预测结果结合,以估计自由临床访谈中患者报告的互动质量。我们在多种标准模型类型(包括Ridge、SVR、MLP、GRU和BiLSTM)上进行评估,所有模型均基于107段精神科患者与访谈者之间的双人对话转录文本提取的句子嵌入训练。结果表明,通过简单平均融合访谈者判断与模型预测可获得最佳整体性能。仅访谈者基线相关系数为0.365;在完全自动模型中,Ridge表现最佳(r = 0.286),BiLSTM为r = 0.270;而结合两者的最强结果为BiLSTM整合模型(r = 0.403)。研究发现,自动语言分析与访谈者判断捕捉了患者体验的不同侧面,其结合能比单一来源更准确地逼近患者自身报告。

原文摘要 · Abstract (English)

Understanding how psychiatric patients subjectively experienced a clinical conversation is important for feedback and alliance-related process monitoring. While interviewers form post-session judgments about patient experience, these judgments do not always match patients' self-reports. Automatic approaches for predicting perceived interaction quality from conversation have been proposed, but it remains unclear whether such approaches can complement human judgment rather than simply replicate it. To address this gap, we evaluate a clinician-support framework in which post-session interviewer ratings are combined with automatic language-based predictions to estimate patient-reported interaction quality in free clinical interviews. We assess this integration across multiple standard model types, including Ridge, SVR, MLP, GRU, and BiLSTM, all trained on sentence embeddings extracted from dyadic transcripts of 107 free conversations between psychiatric patients and interviewers. Our results show that combining interviewer judgments with model predictions through simple averaging yields the strongest overall performance. The interviewer-only baseline reached a Pearson correlation of 0.365. Among fully automatic models, Ridge achieved the strongest Pearson correlation (r = 0.286), while BiLSTM achieved r = 0.270. The strongest result was obtained by BiLSTM interviewer integration (r = 0.403). Our findings suggest that automatic language analysis and interviewer judgment capture complementary aspects of patient experience and that their combination provides a more accurate approximation of the patient's own report than either source alone.

医疗对话语言分析评估系统临床支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。