arXiv:2603.00028cs.CL2026-03被引 3

构建医疗沟通评测基准,评估大模型理解患者医生对话能力。

EPPCMinerBen: A Novel Benchmark for Evaluating Large Language Models on Electronic Patient-Provider Communication via the Patient Portal

  • 基于耶鲁医院1933条消息构建三任务评测集,涵盖意图识别与信息提取。
  • 大模型在证据抽取上表现最佳(Llama-3.1-70B F1=82.84%),小模型在细粒度分类中差距超30%。
  • 适合医疗AI研究者用于评估模型在真实医患沟通中的理解与推理能力。

有效的医疗沟通对治疗效果和依从性至关重要。随着患者-医生交流转向安全消息平台,分析电子患者-医生通信(EPPC)数据既重要又具挑战性。本文提出EPPCMinerBen,一个用于评估大语言模型(LLM)在检测沟通模式和提取信息方面能力的基准。该基准包含三个子任务:代码分类、子代码分类和证据抽取。基于耶鲁纽黑文医院患者门户的752条安全消息中的1,933条专家标注语句,评估模型在识别交流意图和支持性文本方面的能力。基准涵盖多种大模型在零样本与少样本设置下的表现,数据将通过美国国家癌症研究所(NCI)癌症数据服务发布。实验结果表明,模型在不同任务与设置下表现差异显著。Llama-3.1-70B在证据抽取任务中表现最优(F1: 82.84%),而Llama-3.3-70b-Instruct在代码分类中领先(F1: 67.03%)。DeepSeek-R1-Distill-Qwen-32B在子代码分类中表现突出(F1: 48.25%),sdoh-llama-3-70B则表现稳定。较小模型整体表现较差,尤其在子代码分类中性能下降超过30%。少样本提示显著提升多数任务表现。结果表明,大尺寸、指令微调的模型在EPPCMinerBen任务中总体更优,尤其在证据抽取方面;而小模型在细粒度推理任务中存在明显短板。EPPCMinerBen为话语层面的理解提供了基准,支持未来模型泛化与医患沟通分析研究。

原文摘要 · Abstract (English)

Effective communication in health care is critical for treatment outcomes and adherence. With patient-provider exchanges shifting to secure messaging, analyzing electronic patient-communication (EPPC) data is both essential and challenging. We introduce EPPCMinerBen, a benchmark for evaluating LLMs in detecting communication patterns and extracting insights from electronic patient-provider messages. EPPCMinerBen includes three sub-tasks: Code Classification, Subcode Classification, and Evidence Extraction. Using 1,933 expert annotated sentences from 752 secure messages of the patient portal at Yale New Haven Hospital, it evaluates LLMs on identifying communicative intent and supportive text. Benchmarks span various LLMs under zero-shot and few-shot settings, with data to be released via the NCI Cancer Data Service. Model performance varied across tasks and settings. Llama-3.1-70B led in evidence extraction (F1: 82.84%) and performed well in classification. Llama-3.3-70b-Instruct outperformed all models in code classification (F1: 67.03%). DeepSeek-R1-Distill-Qwen-32B excelled in subcode classification (F1: 48.25%), while sdoh-llama-3-70B showed consistent performance. Smaller models underperformed, especially in subcode classification (>30% F1). Few-shot prompting improved most tasks. Our results show that large, instruction-tuned models generally perform better in EPPCMinerBen tasks, particularly evidence extraction while smaller models struggle with fine-grained reasoning. EPPCMinerBen provides a benchmark for discourse-level understanding, supporting future work on model generalization and patient-provider communication analysis. Keywords: Electronic Patient-Provider Communication, Large language models, Data collection, Prompt engineering

医疗AI大模型评测对话理解医学文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。