首个长期喉部病变语音数据集,助力罕见病语音检测研究
RRP-Voice: A Longitudinal Dataset and Benchmark for Recurrent Respiratory Papillomatosis Detection

- 构建26名患者长达十年的语音追踪数据集
- 验证语音特征与喉镜状态相关,而非固定声纹特征
- 适合临床语音分析、罕见病监测方向研究者
深度学习推动了病理语音检测进展,但罕见喉部疾病因数据稀缺仍被忽视。复发性呼吸性乳头状瘤(RRP)即为典型:由人乳头瘤病毒引发,患者多年间反复发作与术后缓解交替。现有横断面数据集无法支持连续语音监测。本文首次发布针对RRP的纵向语音数据集,包含26名患者最长十年随访记录,每轮采集持续元音与句级语句,均由耳鼻喉科医生标注,并同步喉镜确认。基于此资源,建立系统性基准,涵盖手工特征、端到端深度网络、自监督预训练模型及最新音频大模型,均在会话级交叉验证下评估,辅以患者级审计。个体纵向分析进一步证实,横断面判别信号反映喉镜疾病状态,而非稳定说话人特征。本工作为低资源临床场景下的罕见纵向病理语音任务奠定基础。
原文摘要 · Abstract (English)
Deep learning has advanced pathological voice detection rapidly, yet rare laryngeal diseases remain underexplored due to data scarcity. Recurrent Respiratory Papillomatosis (RRP) exemplifies this gap: an HPV-induced disease of the larynx in which patients oscillate between recurrence and post-surgical remission over the years. RRP demands continuous voice monitoring that existing cross-sectional corpora cannot support. We introduce the first longitudinal voice dataset for RRP, comprising recordings from 26 patients with up to ten years of follow-up. Each session pairs sustained vowels with sentence-level utterances, which are annotated by otolaryngologists and confirmed synchronously with laryngoscopy. Building on this resource, we establish a systematic benchmark spanning handcrafted features, end-to-end deep networks, self-supervised pretrained models, and recent audio large language models, all evaluated under session-level cross-validation with patient-level audit. Per-subject longitudinal analyses further confirm that the cross-sectional discriminative signal reflects laryngoscopic disease state rather than stable speaker attributes. This work lays a foundation for rare longitudinal pathological voice tasks in low-resource clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。