arXiv:2506.13970cs.SDcs.AI2025-06被引 2

用AI分析婴儿哭声预测疾病,解决数据少、模型大、跨场景差问题。

Making deep neural networks work for medical audio: representation, compression and domain adaptation

  • 用成人语音数据迁移学习,提升小样本婴儿哭声识别准确率。
  • 通过张量分解压缩循环网络,实现数百倍压缩且保持高精度。
  • 提出适配音频的域自适应方法,让模型更通用,适合临床落地。

本论文针对医疗音频信号机器学习应用中的技术挑战展开研究。肺部、心脏及声音等生理声响蕴含重要健康信息,但目前仍主要依赖医生听诊判断。自动化分析可标准化处理流程,在医疗资源匮乏地区实现筛查,并发现人耳难以察觉的细微模式,助力早期诊断。聚焦婴儿哭声预测疾病,本文在四方面作出贡献:一、在低数据条件下,利用大规模成人语音数据库进行神经迁移学习,显著提升婴儿哭声分析模型的准确性与鲁棒性;二、提出一种端到端的循环网络压缩方法,基于张量分解,无需后期处理,压缩率达数百倍,生成轻量级可部署模型;三、设计专用于音频的域自适应技术,融合计算机视觉方法,缓解数据偏差,增强跨域泛化能力,同时保持原始数据性能;四、联合全球临床专家发布首个公开的婴儿哭声数据集。该工作为将婴儿哭声确立为重要生命体征奠定基础,凸显人工智能驱动音频监测在普惠医疗中的变革潜力。

原文摘要 · Abstract (English)

This thesis addresses the technical challenges of applying machine learning to understand and interpret medical audio signals. The sounds of our lungs, heart, and voice convey vital information about our health. Yet, in contemporary medicine, these sounds are primarily analyzed through auditory interpretation by experts using devices like stethoscopes. Automated analysis offers the potential to standardize the processing of medical sounds, enable screening in low-resource settings where physicians are scarce, and detect subtle patterns that may elude human perception, thereby facilitating early diagnosis and treatment. Focusing on the analysis of infant cry sounds to predict medical conditions, this thesis contributes on four key fronts. First, in low-data settings, we demonstrate that large databases of adult speech can be harnessed through neural transfer learning to develop more accurate and robust models for infant cry analysis. Second, in cost-effective modeling, we introduce an end-to-end model compression approach for recurrent networks using tensor decomposition. Our method requires no post-hoc processing, achieves compression rates of several hundred-fold, and delivers accurate, portable models suitable for resource-constrained devices. Third, we propose novel domain adaptation techniques tailored for audio models and adapt existing methods from computer vision. These approaches address dataset bias and enhance generalization across domains while maintaining strong performance on the original data. Finally, to advance research in this domain, we release a unique, open-source dataset of infant cry sounds, developed in collaboration with clinicians worldwide. This work lays the foundation for recognizing the infant cry as a vital sign and highlights the transformative potential of AI-driven audio monitoring in shaping the future of accessible and affordable healthcare.

医疗音频婴儿哭声模型压缩域自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。