构建临床操作中的多模态对话数据集,助力团队协作研究
CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation
- 从医疗模拟中采集音视频、生理信号与双视角操作数据
- 发现现有大模型在处理真实临床数据时表现显著下降
- 适合医学人工智能、人机协同与医疗团队行为研究者
在临床操作中,团队协作是决定最终结果的关键因素。为理解团队在操作中的协作过程,我们收集了基于医疗模拟的多模态对话数据集 ClinDial。该数据集包含音频及其转录文本、患者假人模拟的生理信号,以及来自两个摄像头视角的团队操作记录。我们采用现有框架对行为进行标注,以分析团队协作流程。研究揭示数据集三大特征:标签不平衡、自然丰富的交互、多模态融合。实验表明,现有大语言模型在处理此类数据时面临显著挑战,凸显开发更适应真实临床数据方法的必要性。代码已开源于 https://github.com/MichiganNLP/CliniDial。
原文摘要 · Abstract (English)
In clinical operations, teamwork can be the crucial factor that determines the final outcome. Prior studies have shown that sufficient collaboration is the key factor that determines the outcome of an operation. To understand how the team practices teamwork during the operation, we collected CliniDial from simulations of medical operations. CliniDial includes the audio data and its transcriptions, the simulated physiology signals of the patient manikins, and how the team operates from two camera angles. We annotate behavior codes following an existing framework to understand the teamwork process for CliniDial. We pinpoint three main characteristics of our dataset, including its label imbalances, rich and natural interactions, and multiple modalities, and conduct experiments to test existing LLMs' capabilities on handling data with these characteristics. Experimental results show that CliniDial poses significant challenges to the existing models, inviting future effort on developing methods that can deal with real-world clinical data. We open-source the codebase at https://github.com/MichiganNLP/CliniDial
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。