构建144小时多模态医疗数据集,助力疾病自动检测与非语言行为分析。
CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions

- 收集612人跨12类疾病及健康对照的视频访谈,覆盖多模态信号
- 包含144小时视频,提供药物、情绪等结构化元数据支持可靠分析
- 适用于精神、神经、呼吸系统疾病的自动诊断与情绪情境建模
自动分析多模态语音在检测和监测多种神经系统、精神及呼吸系统疾病方面展现出巨大潜力。然而,该领域进展受限于现有公开数据集规模小、仅聚焦单一疾病且以语音为主。此外,若教育水平、用药情况、共病或情绪状态等关键混杂因素未充分记录,计算分析的可靠性与可解释性将受到影响。为解决上述问题,我们推出了CARE v1.0,一个经过精心整理的多模态英文数据集,包含约144小时短视频访谈,来自612名个体,涵盖12种医学状况及对照组。每段视频均配有全面的临床相关多模态描述,以及涵盖用药、生活影响和情绪表达等结构化元数据。该数据集的广度与异质性支持多种应用,包括自动疾病与症状检测、情绪刺激情境下语音与非言语行为的多模态建模,以及疾病轨迹与应对过程的研究。
原文摘要 · Abstract (English)
Automatic analysis of multimodal speech has shown strong potential for computationally detecting and monitoring a wide range of neurological, psychiatric, and respiratory conditions. However, progress in this field is limited by existing publicly accessible datasets, which are often small in scale, focused on a single condition or disease, and primarily speech focused. Moreover, if key confounding variables such as education, medication use, comorbidities, or mood state are insufficiently documented, the reliability and interpretability of computational analyses are further compromised. To address these limitations, we introduce CARE v1.0, a curated multimodal English dataset of approximately 144 hours of short video interviews collected from 612 individuals across 12 medical conditions plus a control cohort. For each video, a comprehensive set of clinically relevant multimodal descriptors is provided, alongside structured metadata covering factors such as medication, life impacts, and expressed emotions. The corpus's breadth and heterogeneity support a wide range of applications, including automatic disease and symptom detection, multimodal modelling of speech and non-verbal behaviour under emotionally charged contexts, and studies of disease trajectories and coping processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。