arXiv:2411.06549cs.AIcs.CL2024-11被引 1

用少量真实消息训练大模型,生成符合医疗风格的隐私安全患者消息。

In-Context Learning for Preserving Patient Privacy: A Framework for Synthesizing Realistic Patient Portal Messages

  • 基于少量去标识化消息,通过少样本引导生成真实风格文本。
  • 生成消息在质量和风格上优于现有方法和数据集。
  • 适合医疗数据合成、隐私保护研究及临床工作流优化者使用。

自新冠疫情以来,临床医生收到的患者门户消息数量大幅增加,导致职业倦怠。据我们所知,目前尚无大规模公开的患者门户消息语料库可供研究人员用于优化临床工作流程。基于与区域性医院的持续合作,本研究提出一种基于大语言模型的可配置、逼真的患者门户消息生成框架。该方法采用少样本接地文本生成,仅需少量去标识化的患者消息即可帮助大模型更准确地匹配真实数据的风格与语气。团队中的临床专家认为该框架符合HIPAA要求,而现有合成文本生成方法无法确保所有敏感信息均被保护。通过广泛的定量与人工评估,结果表明该框架生成的数据质量高于现有生成方法及所有相关数据集。我们认为此工作为(i)发布与真实样本风格一致的大规模合成患者消息数据集,以及(ii)实现低人工成本的合规隐私保护数据生成提供了可行路径。

原文摘要 · Abstract (English)

Since the COVID-19 pandemic, clinicians have seen a large and sustained influx in patient portal messages, significantly contributing to clinician burnout. To the best of our knowledge, there are no large-scale public patient portal messages corpora researchers can use to build tools to optimize clinician portal workflows. Informed by our ongoing work with a regional hospital, this study introduces an LLM-powered framework for configurable and realistic patient portal message generation. Our approach leverages few-shot grounded text generation, requiring only a small number of de-identified patient portal messages to help LLMs better match the true style and tone of real data. Clinical experts in our team deem this framework as HIPAA-friendly, unlike existing privacy-preserving approaches to synthetic text generation which cannot guarantee all sensitive attributes will be protected. Through extensive quantitative and human evaluation, we show that our framework produces data of higher quality than comparable generation methods as well as all related datasets. We believe this work provides a path forward for (i) the release of large-scale synthetic patient message datasets that are stylistically similar to ground-truth samples and (ii) HIPAA-friendly data generation which requires minimal human de-identification efforts.

医疗生成隐私保护大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。