用可控生成提升呼吸音分类数据质量,解决小样本与噪声问题。
C2GA: A Class-Controllable Generative Augmentation Framework for Respiratory Sound Classification

- 通过条件VQ-VAE构建解耦的语义离散潜空间,分离声音特征与类别原型。
- 利用Transformer生成符合标签的声学令牌序列,合成高保真梅尔频谱图。
- 适合医疗音频分析、数据稀缺场景下的模型增强,提升分类鲁棒性。
呼吸音分类在肺部疾病临床诊断中至关重要,但受限于真实听诊数据集规模小、噪声严重及类别不平衡。传统音频增强方法易扭曲细微病理特征,现有基于变分自编码器(VAE)或生成对抗网络(GAN)的生成方法在样本保真度和类别语义控制方面表现不足,尤其在监督信息稀缺时。为此,本文提出C2GA框架:首先使用条件向量量化变分自编码器(VQ-VAE)构建语义丰富的离散潜空间,将局部声学标记与全局类别原型显式解耦;随后训练基于Transformer的自回归先验,生成标签一致的令牌序列;最后将生成令牌与对应类别原型融合并解码为高质量梅尔频谱图,用于数据增强。结果表明,C2GA提供了一种有效且语义可靠的呼吸音增强策略,通过可控且高保真的数据生成,显著提升了呼吸音分类在真实临床场景中的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Background: Respiratory sound classification plays a critical role in the clinical identification of pulmonary pathologies. However, its performance is often hindered by the limited size, severe noise, and class imbalance of real-world auscultation datasets. Although conventional audio augmentation techniques are easy to implement, they may inadvertently distort subtle pathological characteristics. Meanwhile, existing Variational Autoencoder (VAE)- or Generative Adversarial Network (GAN)-based generative approaches often suffer from limited sample fidelity and insufficient controllability over class semantics, particularly under conditions of scarce supervision. Methods: To overcome these limitations, we propose C2GA, a class-controllable generative augmentation framework. C2GA first constructs a semantically rich discrete latent space using a conditional Vector-Quantized Variational Autoencoder (VQ-VAE), in which local acoustic tokens are explicitly decoupled from global class prototypes. Subsequently, a Transformer-based autoregressive prior is trained to generate label-consistent token sequences. These generated tokens are then fused with the corresponding class prototypes and decoded into high-fidelity Mel-spectrograms for data augmentation. Conclusion: These results indicate that C2GA provides an effective and semantically reliable augmentation strategy for respiratory sound analysis. By enabling controllable and high-quality data generation, the proposed framework offers a promising solution for improving the robustness and generalization of respiratory sound classification in realistic clinical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。