用语音检测喉癌,36个模型公开可复现。
A Classification Benchmark for Artificial Intelligence Detection of Laryngeal Cancer from Patient Voice
- 构建36个模型,基于语音和临床数据分类喉部病变。
- 最佳模型准确率达83.7%,AUROC达91.8%。
- 开源所有模型与数据,推动非侵入式癌症筛查研究。
喉癌病例预计未来将显著增加,当前诊断路径效率低下,给患者和医疗系统带来压力。人工智能可通过分析患者语音实现喉癌的非侵入式检测,有助于更高效地筛选转诊。本研究解决该领域缺乏可复现方法的问题,提出一个基准套件,包含36个在开源数据集上训练和评估的模型,用于区分良性与恶性嗓音病理。所有模型均开源,支持未来研究。我们评估了三种算法与三种音频特征集,包括仅音频输入及融合人口学和症状数据的多模态输入。最佳模型达到平衡准确率83.7%、敏感度84.0%、特异度83.3%、AUROC 91.8%。
原文摘要 · Abstract (English)
Cases of laryngeal cancer are predicted to rise significantly in the coming years. Current diagnostic pathways are inefficient, putting undue stress on both patients and the medical system. Artificial intelligence offers a promising solution by enabling non-invasive detection of laryngeal cancer from patient voice, which could help prioritise referrals more effectively. A major barrier in this field is the lack of reproducible methods. Our work addresses this challenge by introducing a benchmark suite comprising 36 models trained and evaluated on open-source datasets. These models classify patients with benign and malignant voice pathologies. All models are accessible in a public repository, providing a foundation for future research. We evaluate three algorithms and three audio feature sets, including both audio-only inputs and multimodal inputs incorporating demographic and symptom data. Our best model achieves a balanced accuracy of 83.7%, sensitivity of 84.0%, specificity of 83.3%, and AUROC of 91.8%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。