用患者健康记录提升AI问诊准确率,让个人医疗数据真正有用。
Evaluating the Utility of Personal Health Records in Personalized Health AI

- 用大模型分析患者健康记录,比无上下文回答更精准。
- 加入完整病历后,回答帮助性显著提升(p<0.001)。
- 发现模型对时间错乱、虚构信息理解不足,可针对性改进。
患者管理的个人健康记录(PHRs)有望增强患者对自身健康的理解;但记录内容复杂,可能阻碍洞察。本研究评估大语言模型(Gemini 3.0 Flash)在提供临床数据上下文时,对用户健康问题的回答帮助性。共收集2,257个用户问题,来自三种来源:简短网络搜索、聊天机器人模板生成的长问题、以及患者向医护团队提问的真实电话记录。问题匹配去标识化的PHRs(来自1,945份记录)。生成三种情境下的回答:(1) 无PHR上下文;(2) 基础摘要(人口统计、疾病、用药);(3) 完整临床笔记。使用现有评分框架(SHARP)和新开发的错误模式框架进行评估。全量由自动评分器评估,子集(n=95)由临床医生评分,两者均知晓完整PHR上下文。结果显示,所有问题类型中加入PHR数据后回答帮助性显著提升(p < 0.001,配对t检验)。同时观察到安全性、准确性、相关性和个性化均有潜在提升。新框架识别出模型在理解复杂PHR中的关键缺陷,如时间错乱、罕见但重要的虚构内容。结果表明PHR数据能有效支持多样用户需求,并提供监测模型短板的评估框架。研究推动进一步探索用户从理解健康记录中获得的实际益处。
原文摘要 · Abstract (English)
Patient-managed Personal Health Records (PHRs) promises to empower patients to better understand their health; but information in the record is complex, potentially hindering insights. In this study, we assess the potential of large language models (LLMs, Gemini 3.0 Flash) to provide helpful answers to user health queries, when provided clinical data from PHRs as context. A total of 2,257 user queries were drawn from 3 different distributions to represent patient questions: shorter web search queries, longer questions derived from templates of chatbot conversations, and questions patients asked to their healthcare team (patient calls). Queries were matched with de-identified PHRs (from a pool of 1,945). Gemini responses were generated (1) without PHR context; (2) with a basic summary of demographics, conditions, and medications; (3) with full, extensive clinical notes. For evaluation, we leveraged an existing rating framework (SHARP), and developed a new framework for specific error modes when interpreting PHRs. Evaluation was performed using autoraters for the full set, and with clinician ratings for a subset (n=95), with both sets of raters knowing the full PHR context. We see significant improvements in the helpfulness of answers to all question types with PHR data (p < 0.001, paired t-test). We also observe potential gains in safety, accuracy, relevance and personalization of answers. Our PHR evaluation framework further identifies gaps in LLM understanding of particular aspects of complex PHRs, such as temporal disorientation, and rare but meaningful confabulations. These results suggest potential for PHR data to help people with a wide range of user needs; and provide a framework for monitoring for gaps in LLM answers based on PHR context. This study motivates further work to assess and realize potential benefits to users from understanding their health records.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。