首次系统评估27个语音生物标志物数据集的可发现性与可用性,助力精神与神经退行性疾病诊断。
Systematic FAIRness Assessment of Open Voice Biomarker Datasets for Mental Health and Neurodegenerative Diseases
- 基于FAIR原则对27个公开语音数据集进行多维度评分
- 发现可发现性高但可访问性、互操作性差,精神健康类数据差异更大
- 建议采用领域专用元数据和合规存储库,提升数据重用价值
语音生物标志物——如言语、咳嗽、呼吸等人类发声信号——是实现精神健康与神经退行性疾病大规模、无创检测与监测的有前景工具。然而,其临床应用受限于公开数据集质量参差不齐、可用性差。为填补这一空白,我们首次对27个聚焦此类疾病的公开语音生物标志物数据集进行了系统性FAIR(可发现、可访问、可互操作、可重用)评估。采用FAIR数据成熟度模型及加权评分方法,从子原则、原则到综合层面进行评价。分析显示,所有数据集在可发现性方面表现良好,但在可访问性、互操作性和可重用性上存在显著差异与薄弱环节。精神健康类数据集的FAIR得分波动更大,而神经退行性疾病数据集相对更一致。数据存储库的选择也显著影响评分结果。为提升数据质量和临床应用潜力,我们建议采用结构化、领域特定的元数据标准,优先选择符合FAIR规范的存储平台,并定期应用结构化评估框架。这些发现为提升数据互操作性与重用性提供了可操作指引,加速语音生物标志物技术的临床转化。
原文摘要 · Abstract (English)
Voice biomarkers--human-generated acoustic signals such as speech, coughing, and breathing--are promising tools for scalable, non-invasive detection and monitoring of mental health and neurodegenerative diseases. Yet, their clinical adoption remains constrained by inconsistent quality and limited usability of publicly available datasets. To address this gap, we present the first systematic FAIR (Findable, Accessible, Interoperable, Reusable) evaluation of 27 publicly available voice biomarker datasets focused on these disease areas. Using the FAIR Data Maturity Model and a structured, priority-weighted scoring method, we assessed FAIRness at subprinciple, principle, and composite levels. Our analysis revealed consistently high Findability but substantial variability and weaknesses in Accessibility, Interoperability, and Reusability. Mental health datasets exhibited greater variability in FAIR scores, while neurodegenerative datasets were slightly more consistent. Repository choice also significantly influenced FAIRness scores. To enhance dataset quality and clinical utility, we recommend adopting structured, domain-specific metadata standards, prioritizing FAIR-compliant repositories, and routinely applying structured FAIR evaluation frameworks. These findings provide actionable guidance to improve dataset interoperability and reuse, thereby accelerating the clinical translation of voice biomarker technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。