arXiv:2510.03997cs.CL2025-10中稿 · npj Digital Medici…被引 1

用大模型分析百万医生评价,绘制患者眼中的医生特质地图。

Mapping Patient-Perceived Physician Traits from Nationwide Online Reviews with LLMs

  • 用大模型从评论中提取十类医生特质得分
  • 男性医生整体评分更高,外科医生人际能力更强
  • 识别出四种医生类型,适合研究医疗公平与患者选择

了解患者如何感知医生对提升信任、沟通和满意度至关重要。患者越来越多地使用大语言模型(LLMs)来总结医生评价并影响就医选择,但全国范围内的患者感知医生特质仍缺乏系统刻画。我们提出一种基于大模型的分析流程,从评论文本中提取十类患者感知医生特质得分:五类大五型特质和五类患者导向维度。基于美国100万位医生的410万条评论(涉及226,999名医生),我们通过多模型对比和人工专家基准验证了该流程的有效性。大模型与人工评分高度一致,且特质分与评分相关但仍有独立变异。两个全国性模式显现:男性医生在所有特质上得分更高,临床能力差距最大;专科差异由就诊情境驱动,外科专科在人际能力上领先,精神科最低。聚类分析识别出四种医生原型,包括‘全面高分’(33.8%)、‘全面低分’(22.6%)。该大模型生成的医生特质图谱揭示了大模型如何解读美国临床队伍,为未来关于公平性、偏见及大模型辅助医生选择的研究提供基础。

原文摘要 · Abstract (English)

Understanding how patients perceive their physicians is essential to improving trust, communication, and satisfaction. Patients increasingly consult large language models (LLMs) to summarize physician reviews and shape provider choices, yet the national landscape of patient-perceived physician traits remains poorly characterized. We present an LLM-based pipeline that extracts ten patient-perceived physician trait scores from review text: five Big-Five-style and five patient-oriented dimensions. From one million U.S. physicians, we analyze 4.1 million reviews of 226,999 physicians. We validate the pipeline through multi-model comparison and human expert benchmarking. LLM and human-rater trait scores from reviews are consistent. Trait scores correlate strongly with review rating scores yet retain substantial independent variance. Two national-scale patterns emerge: male physicians receive higher trait scores across all traits, with the largest gap in clinical competence; specialty differences are driven by encounter context, with surgical specialties leading interpersonal qualities and psychiatry lowest. Cluster analysis identifies four physician archetypes, from "Uniform High" (33.8%, high across traits) to "Uniform Low" (22.6%, low across traits). This map of LLM-derived physician traits exposes how LLMs read the U.S. clinical workforce. Pending clinical validation, it opens future research on fairness, bias, and how LLM-mediated provider search shapes patient choice.

大模型医生评价患者感知医疗公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。