构建多模态医患对话数据集,助力AI理解真实临床互动。
Longitudinal and Multimodal Recording System to Capture Real-World Patient-Clinician Conversations for AI and Encounter Research: Protocol
- 用360度音视频记录医患面对面交流,融合问卷与病历数据。
- 97%医生、75%患者同意参与,96%完成问卷,76%录音完整。
- 为医疗AI研究提供可复现的多模态数据采集框架,适合临床研究者。
AI在医疗中的潜力依赖于反映患者与医生关注点的数据。现有模型多基于电子健康记录(EHR),仅包含生物指标,缺乏医患互动信息。而诊疗核心正是语音、文字与视频交织的交流过程。为此,我们设计并评估了一套纵向、多模态系统,用于在梅奥诊所内分泌科门诊采集患者-医生互动数据。通过360度摄像头录制,结合访后问卷(共情、满意度、节奏、治疗负担)与EHR提取人口学及临床数据。可行性通过五个指标评估:医生同意率、患者同意率、录音成功率、问卷完成率、跨模态数据关联率。2025年1月启动招募,至8月,36名合格医生中有35名(97%)同意,281名患者中212名(75%)同意参与。已同意的就诊中,162次(76%)完成完整录音,204次(96%)完成问卷。本研究旨在验证该框架的可行性,为长期多模态数据集建设提供可复制的模板,推动涵盖诊疗复杂性的医疗AI发展。
原文摘要 · Abstract (English)
The promise of AI in medicine depends on learning from data that reflect what matters to patients and clinicians. Most existing models are trained on electronic health records (EHRs), which capture biological measures but rarely patient-clinician interactions. These relationships, central to care, unfold across voice, text, and video, yet remain absent from datasets. As a result, AI systems trained solely on EHRs risk perpetuating a narrow biomedical view of medicine and overlooking the lived exchanges that define clinical encounters. Our objective is to design, implement, and evaluate the feasibility of a longitudinal, multimodal system for capturing patient-clinician encounters, linking 360 degree video/audio recordings with surveys and EHR data to create a dataset for AI research. This single site study is in an academic outpatient endocrinology clinic at Mayo Clinic. Adult patients with in-person visits to participating clinicians are invited to enroll. Encounters are recorded with a 360 degree video camera. After each visit, patients complete a survey on empathy, satisfaction, pace, and treatment burden. Demographic and clinical data are extracted from the EHR. Feasibility is assessed using five endpoints: clinician consent, patient consent, recording success, survey completion, and data linkage across modalities. Recruitment began in January 2025. By August 2025, 35 of 36 eligible clinicians (97%) and 212 of 281 approached patients (75%) had consented. Of consented encounters, 162 (76%) had complete recordings and 204 (96%) completed the survey. This study aims to demonstrate the feasibility of a replicable framework for capturing the multimodal dynamics of patient-clinician encounters. By detailing workflows, endpoints, and ethical safeguards, it provides a template for longitudinal datasets and lays the foundation for AI models that incorporate the complexity of care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。