arXiv:2607.10596eess.AS2026-07

ECHOv2通过分频段学习与跨频交互,提升机器异常声音检测的精度。

ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection

论文配图:ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection
图 1 · 摘自论文原文
  • 将音频分频段学习,捕捉细微频谱特征
  • 引入两级自蒸馏与跨频监督,增强频率间关联建模
  • 支持多种设备和噪声环境,适合工业场景应用

机器异常声音检测(ASD)需要在有限监督下具备捕捉声音微小偏差的鲁棒音频表征能力。现有预训练音频主干模型未能充分捕捉机器声音的频段特性。为此,我们提出ECHOv2,一种分频段模型,通过学习局部频段内表征来捕捉细粒度频谱模式,并结合两级自蒸馏策略与显式跨频监督,建模跨频依赖关系。跨频分支实现全局上下文对齐与掩码子频段重建,引入多个汇总标记以可控粒度进行结构化聚合,实现训练中子频段间的区域感知交互。该设计使ECHOv2在多样化机器类型和噪声运行条件下保持稳定表征质量。为实现预训练音频主干的公平一致评估,我们在DCASE 2020-2025上建立统一的ASD基准,包含两种互补协议:基于嵌入的评估(冻结表征判别性)与基于适配的评估(下游可迁移性)。消融实验验证了频段内学习、跨频监督及结构化聚合粒度对鲁棒ASD表征学习的有效性。结果表明,结构化跨频建模为ASD表征学习提供了强大且可扩展的框架,可作为未来研究基础。模型与基准已开源:https://github.com/yucongzh/ECHOv2 与 https://github.com/yucongzh/ASD_Benchmark。

原文摘要 · Abstract (English)

Machine anomalous sound detection (ASD) requires robust audio representations capable of capturing subtle deviations in machine sounds under limited supervision. Existing pre-trained audio backbones do not fully capture frequency-specific characteristics of machine sounds. To address this, we propose ECHOv2, a band-splitting model that learns localized intra-band representations to capture fine-grained spectral patterns while also incorporating a two-level self-distillation strategy with explicit inter-band supervision to model cross-frequency dependencies. The inter-band branch performs global context alignment and masked sub-band reconstruction, and multiple summary tokens are introduced for structured aggregation with controllable frequency granularity, enabling region-aware interaction across sub-bands during training. This design allows ECHOv2 to robustly handle diverse machine types and noisy operating conditions while maintaining stable representation quality. To enable fair and consistent evaluation of pre-trained audio backbones, we establish a unified ASD benchmark over DCASE 2020-2025 with two complementary protocols: embedding-based evaluation for frozen representation discriminability and adaptation-based evaluation for downstream transferability. Ablation studies confirm the effectiveness of intra-band learning, inter-band supervision, and structured aggregation granularity for robust ASD representation learning. These findings demonstrate that structured cross-band modeling provides a powerful and adaptable framework for ASD representation learning and can serve as a strong foundation for future research. The model and benchmark are fully open-sourced at https://github.com/yucongzh/ECHOv2 and https://github.com/yucongzh/ASD_Benchmark to promote reproducible research.

异常检测音频表征分频段学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。