arXiv:2605.02700eess.AS2026-05

用动态时间建模提升声音障碍的日常检测精度

Neck-Learn: Attention-Based Multiple Instance Learning and Ensemble Framework for Ecological Momentary Assessment

  • 结合梯度提升树与卷积MIL,保留每日语音动态特征
  • 在测试集上对PVH和NPVH的AUC分别达0.879和0.848
  • 适合关注临床语音分析与可穿戴设备监测的研究者

声带功能亢进(VH)是一种常见嗓音障碍,尽管有大量日常语音数据,其便携式检测仍具挑战。现有方法将一周内的颈部加速度计数据压缩为固定长度的个体特征向量,丢失了日间时间动态所蕴含的精细发声特征交互信息。本文提出一种新型混合架构:在日级别分布特征上使用梯度提升树,并结合基于卷积神经网络的多实例学习(MIL)框架,以保留并学习每日内部的时间动态。在独立测试集上,模型超越基准表现(PVH AUC: 0.82,NPVH AUC: 0.77),达到PVH AUC 0.879(排名5)、NPVH AUC 0.848(排名3),同时提供关于两类病理的临床相关洞察。

原文摘要 · Abstract (English)

Vocal hyperfunction (VH) is a prevalent voice disorder whose ambulatory detection remains challenging despite extensive daily voice data. Prior approaches capture week-long neck-surface accelerometer recordings but collapse them into fixed-length subject-level feature vectors, discarding within-day temporal dynamics encoding nuanced voicing feature interactions. We introduce a novel hybrid architecture combining gradient-boosted trees on day-level distributional features with a CNN-based multiple instance learning (MIL) framework that preserves and learns from from temporal dynamics throughout each day. On the held-out test set, our model exceeds the challenge baselines (AUC: 0.82 PVH, 0.77 NPVH), achieving AUCs of 0.879 for PVH (Rank 5) and 0.848 for NPVH (Rank 3), while also providing insights into clinically relevant information about both pathologies.

语音检测多实例学习可穿戴监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。