用MFCC提升哮喘与慢阻肺鉴别准确率,自适应窗口优化时序数据。
Optimizing 2D Input Representations and Sub-phase Fusion Strategies for Differential Diagnosis of Asthma and COPD Using CNN- and GRU-Based Networks

- 采用自适应窗口解决呼吸周期不一致问题,优化时频表示参数。
- 13维MFCC配64点分辨率在循环级达0.877最佳F1分数。
- 复杂融合策略无提升,数据增强反而降低性能,强调真实数据重要性。
本研究对比了VAR模型、梅尔频率倒谱系数(MFCC)矩阵和对数梅尔语谱图在肺音分类中的表现。由于呼吸周期长度不一,传统语谱图存在时序维度不一致问题,本文提出自适应长度窗来固定其时间维度,并通过参数测试优化了频域与时域表示。采用多种卷积神经网络(CNN)架构从子相位的二维表示中提取特征,再通过直接拼接、门控循环单元(GRU)及带注意力机制的GRU等策略融合特征。模型评估基于呼吸周期和受试者两个层面,涵盖多个呼吸周期。研究还探讨了多种数据增强技术以应对数据量不足的问题。最优循环级F1得分(0.877)来自13维MFCC、每子相位64点时间分辨率,配合直接特征拼接;最优受试者级F1得分(0.855)则来自13维MFCC、全周期256点时间分辨率,均使用自适应窗口。数据增强整体降低模型性能,其中mixup效果最好。实验表明,MFCC优于对数梅尔语谱图和VAR模型,而复杂融合策略未带来提升,进一步凸显真实数据在肺音研究中的关键作用。
原文摘要 · Abstract (English)
This study aims to explore the performance of the VAR model in comparison with mel-frequency cepstral coefficient (MFCC) matrices and log-mel spectrograms using deep learning. In pulmonary sound classification, spectrogram-based representations suffer from inconsistent temporal dimensions due to varying respiratory cycle durations. Along with traditional trimming/zero-padding, adaptive-length windowing was presented to fix their temporal dimensions. Their spectral and temporal dimensions were optimized by testing a range of parameters. Different convolutional neural network (CNN) architectures were employed to extract features from the two-dimensional representations obtained over the sub-phases. The extracted sub-phase features were then fused using various strategies including direct concatenation, gated recurrent unit (GRU) network and GRU with attention mechanism. Model performances were assessed through respiratory cycle-based evaluation and subject-based evaluation comprising multiple respiratory cycles. Several data augmentation techniques were also studied to cope with limitations in data size. The best cycle-based F1-score (0.877) was obtained using the MFCC matrices with thirteen coefficients and 64-point time resolution per sub-phase representation followed by direct feature concatenation, and the best subject-based F1-score (0.855) was obtained using the MFCC matrices with thirteen coefficients and 256-point time resolution per full-cycle representation, both obtained by adaptive-length windowing. Augmentation degraded the performance of models overall, yet mixup augmentation was the best among the methods tested. MFCC outperformed log-mel spectrogram and VAR model in differentiation of asthma and COPD. Sophisticated fusion strategies did not improve the diagnosis. Augmentation did not contribute, demonstrating the significance of authentic data in pulmonary sound studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。