arXiv:2602.21165cs.CLcs.AI2026-02被引 2

PVminer能自动识别患者文本中的声音与社会健康因素,提升医疗沟通分析效率。

PVminer: A Domain-Specific Tool to Detect the Patient Voice in Patient Generated Data

  • 构建多标签多分类框架,融合患者特异性BERT与主题模型增强语义理解。
  • 在代码、子代码和组合标签任务上分别取得82.25%、80.14%和77.87%的F1分数。
  • 适合医疗数据研究者与临床语言分析团队使用,支持开源共享。

患者生成的文本(如安全消息、调查问卷和访谈)蕴含丰富的患者声音(PV),反映沟通行为与社会健康决定因素(SDoH)。传统定性编码耗时且难以扩展至大规模跨机构数据。现有机器学习与自然语言处理方法虽提供部分解决方案,但常将以患者为中心的沟通(PCC)与SDoH视为独立任务,或依赖不适用于患者语言的模型。本文提出PVminer,一个针对医疗领域优化的NLP框架,用于结构化患者与医生间安全通信中的患者声音。该框架将患者声音检测建模为多标签多分类任务,结合患者特异性BERT编码器(PV-BERT-base与PV-BERT-large)、无监督主题建模(PV-Topic-BERT)进行主题增强,并采用微调分类器预测代码、子代码与组合层级标签。主题表示在微调与推理阶段被引入以丰富语义输入。PVminer在多层次任务中表现优异,超越生物医学与临床预训练基线,在代码、子代码与组合标签任务上的F1得分分别为82.25%、80.14%与77.87%。消融实验表明,作者身份与主题增强均带来显著性能提升。预训练模型、源代码与文档将公开发布,标注数据集可应研究请求提供。

原文摘要 · Abstract (English)

Patient-generated text such as secure messages, surveys, and interviews contains rich expressions of the patient voice (PV), reflecting communicative behaviors and social determinants of health (SDoH). Traditional qualitative coding frameworks are labor intensive and do not scale to large volumes of patient-authored messages across health systems. Existing machine learning (ML) and natural language processing (NLP) approaches provide partial solutions but often treat patient-centered communication (PCC) and SDoH as separate tasks or rely on models not well suited to patient-facing language. We introduce PVminer, a domain-adapted NLP framework for structuring patient voice in secure patient-provider communication. PVminer formulates PV detection as a multi-label, multi-class prediction task integrating patient-specific BERT encoders (PV-BERT-base and PV-BERT-large), unsupervised topic modeling for thematic augmentation (PV-Topic-BERT), and fine-tuned classifiers for Code, Subcode, and Combo-level labels. Topic representations are incorporated during fine-tuning and inference to enrich semantic inputs. PVminer achieves strong performance across hierarchical tasks and outperforms biomedical and clinical pre-trained baselines, achieving F1 scores of 82.25% (Code), 80.14% (Subcode), and up to 77.87% (Combo). An ablation study further shows that author identity and topic-based augmentation each contribute meaningful gains. Pre-trained models, source code, and documentation will be publicly released, with annotated datasets available upon request for research use.

自然语言处理患者声音医疗文本分析主题建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。