arXiv:2606.15436cs.LGcs.AI2026-06中稿 · ICML

首个咳嗽音频回归基准,评估模型预测年龄等连续健康指标的能力。

Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models

论文配图:Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models
图 1 · 摘自论文原文
  • 构建多模型多目标咳嗽回归基准,测试五种基础模型在六项任务上的表现。
  • 小型MLP头超越均值基线,大模型在大数据集上性能更优,体现数据量与模型容量权衡。
  • 生成式预训练模型(OPERA-GT)在年龄预测中优于判别式模型,适合小样本场景。

呼吸声学基础模型(FMs)在咳嗽分类上表现优异,但其从咳嗽音频预测连续健康指标(如年龄、体重指数、疾病概率)的能力尚未充分探索,而这类被动估计在缺乏体测条件的临床场景中具有重要价值。本文提出一个多模型、多目标咳嗽回归基准,评估五种模型(OPERA-CT、OPERA-CE、OPERA-GT、HeAR、M2D+Resp)在三个数据集上六项任务的表现,采用受试者不重叠协议,并对比线性、小型MLP和全量MLP回归头。小型MLP在30个模型×任务组合中的23个上优于均值基线和线性微调;全量MLP在小临床数据集上过拟合,但在大集合上恢复性能,揭示数据量与头部容量间的权衡。HeAR在Coswara数据集上实现9.12年平均绝对误差(MAE)的最佳年龄预测,其在CIDRZ上的结果因可能的预训练重叠被排除于主要结论外。OPERA-GT在所有三个数据集的年龄预测中优于OPERA-CT,CIDRZ上的差异在种子方差范围内,表明生成式预训练优势可从呼吸扩展至咳嗽。HeAR与M2D+Resp在仅50样本时即接近全性能,而OPERA模型需400样本。跨数据集迁移呈强不对称:大规模多样数据可泛化至小临床人群(CoughVID→CIDRZ:-0.17年),反之则显著退化(CIDRZ→Coswara:+2.43年,+26.6%)。

原文摘要 · Abstract (English)

Respiratory acoustic foundation models (FMs) excel at cough classification, yet their ability to predict continuous health quantities from cough audio remains largely unexplored, despite the clinical value of passive age, BMI, and disease probability estimation in settings where physical measurements are unavailable. We introduce the multi-model, multi-target cough regression benchmark evaluating five FMs (OPERA-CT, OPERA-CE, OPERA-GT, HeAR, M2D+Resp) across six targets on three datasets under subject-disjoint protocols, comparing linear, MLP-small, and full MLP regression heads. MLP-small beats the mean-predictor baseline on all tasks and linear probing in 23 of 30 model x task cases, with full MLP overfitting on small clinical data but recovering on larger sets, revealing a dataset size x head-capacity trade-off. HeAR leads within-dataset age regression on Coswara (9.12 yr MAE); its CIDRZ result is excluded from headline claims owing to possible HeAR-CIDRZ pretraining overlap. OPERA-GT is favored over OPERA-CT on age in all three datasets, with the CIDRZ margin within seed variance, extending a generative-pretraining advantage from breath to cough. HeAR and M2D+Resp reach near-full performance at N = 50 samples while OPERA models require N = 400. Cross-dataset transfer is strongly asymmetric as large diverse data generalises to small clinical populations (CoughVID to CIDRZ: -0.17 yr) but not vice versa (CIDRZ to Coswara: +2.43 yr, +26.6%).

语音分析回归任务医疗AI基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。