arXiv:2606.25116eess.AScs.AI2026-06中稿 · KDD

在可穿戴设备条件下评估呼吸声大模型,发现性能显著下降

BCoughBench: Benchmarking Respiratory Acoustic Foundation Models Under Body-Coupled Wearable Sensor Conditions

论文配图:BCoughBench: Benchmarking Respiratory Acoustic Foundation Models Under Body-Coupled Wearable Sensor Conditions
图 1 · 摘自论文原文
  • 构建模拟可穿戴传感器的评测基准BCoughBench
  • 多数模型在可穿戴场景下准确率下降超9.6%,临床敏感性不达标
  • 适合关注医疗语音模型落地的开发者与临床研究者

呼吸声大模型(FMs)此前仅在手机录音上进行评测,但临床部署正转向体表可穿戴设备,其传感器受组织和骨骼衰减高频信号,导致模型可靠性未知。本文提出BCoughBench,在五个经EBEN模拟的体耦合(BC)传感器条件下,对五种大模型(OPERA-CT/CE/GT、HeAR、M2D+Resp)在五个标注咳嗽数据集上进行九项分类任务(AUROC、95%特异性下的敏感性、期望校准误差)和三项年龄回归任务(与均值预测基线比较的平均绝对误差)。结果显示,平均AUROC从手机条件下的0.785降至0.689–0.723,降幅最大为颞部振动采集(Δ = -0.096),最小为软入耳式(Δ = -0.062)。所有模型在多数疾病任务中均未达到临床敏感性阈值(Se@Sp95 ≥ 0.20)。性别分类在CIDRZ队列中性能崩溃(AUROC 0.954 → 0.596–0.628,Δ = -0.341),而新冠检测几乎不受影响(Δ = -0.004)。年龄回归表现稳健,于CoughVID数据集上额头加速度计使误差从9.61年降至8.97年;HeAR在回归与人口统计任务中领先,M2D+Resp在疾病与特征任务中表现更优。该基准提供可复现的可穿戴环境下大模型评测框架。

原文摘要 · Abstract (English)

Respiratory acoustic foundation models (FMs) are benchmarked exclusively on smartphone recordings, yet clinical deployment increasingly targets body-coupled (BC) wearables whose sensors attenuate high-frequency content through tissue and bone, leaving FM reliability uncharacterised. We introduce BCoughBench, evaluating five FMs (OPERA-CT/CE/GT, HeAR, M2D+Resp) on nine classification tasks (AUROC, sensitivity at 95% specificity, Expected Calibration Error) and three age regression tasks (MAE vs. a mean-predictor baseline) across five EBEN-simulated BC sensor conditions on five labeled cough datasets. Mean AUROC declines from 0.785 (smartphone) to 0.689-0.723, degrading most under temple vibration pickup ($Δ$ = -0.096) and least under the soft in-ear ($Δ$ = -0.062). No FM meets the clinical sensitivity threshold (Se@Sp95 $\geq$ 0.20) on most disease tasks under any BC sensor. Sex classification on the CIDRZ cohort collapses (AUROC 0.954 to 0.596-0.628, $Δ$ = -0.341) while COVID detection is nearly unaffected ($Δ$ = -0.004). Age regression is robust, improving under the forehead accelerometer on CoughVID (MAE 9.61 to 8.97 yr); HeAR leads on regression and demographic tasks, M2D+Resp on disease and characteristic tasks. BCoughBench provides a reproducible framework for FM evaluation under wearable conditions.

大模型评测可穿戴设备呼吸分析语音诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。