arXiv:2601.13513cs.SDeess.AS2026-01中稿 · ICASSP 2026

用物理模型修复损坏麦克风信号,提升声源分类准确率

Event Classification by Physics-informed Inpainting for Distributed Multichannel Acoustic Sensor with Partially Degraded Channels

  • 基于逆时迁移的无学习修复方法,重建缺失信号
  • 右角布局下准确率提升13.1点(9.7%→22.8%)
  • 适合传感器布局未知或部分失效场景

分布式多通道声学传感(DMAS)可实现大范围声事件分类(SEC),但当多通道信号受损且测试时传感器布局与训练不一致时性能下降。本文提出一种无需学习、基于物理的反向时间迁移(RTM)信号修复前端:将多通道频谱图在三维网格上使用解析格林函数反向传播形成场景一致图像,再正向投影重建缺失信号,随后提取对数梅尔特征并输入Transformer分类。在包含50个传感器和三种布局(圆形、线性、直角)的ESC-50数据集上评估,单通道信噪比(SNR)在-30至0 dB间采样。相比AST基线、尺度稀疏最大值通道选择和通道交换增强,所提方法在所有布局下均达到最佳或相当性能,尤其在直角布局下准确率提升13.1个百分点(从9.7%到22.8%)。相关性分析显示,空间权重与信噪比的相关性强于与通道-源距离的相关性,且信噪比-权重相关性越高,分类准确率也越高。结果表明,在布局开放和严重通道退化条件下,基于物理的‘重建-投影’预处理能有效补充纯学习方法。

原文摘要 · Abstract (English)

Distributed multichannel acoustic sensing (DMAS) enables large-scale sound event classification (SEC), but performance drops when many channels are degraded and when sensor layouts at test time differ from training layouts. We propose a learning-free, physics-informed inpainting frontend based on reverse time migration (RTM). In this approach, observed multichannel spectrograms are first back-propagated on a 3D grid using an analytic Green's function to form a scene-consistent image, and then forward-projected to reconstruct inpainted signals before log-mel feature extraction and Transformer-based classification. We evaluate the method on ESC-50 with 50 sensors and three layouts (circular, linear, right-angle), where per-channel SNRs are sampled from -30 to 0 dB. Compared with an AST baseline, scaling-sparsemax channel selection, and channel-swap augmentation, the proposed RTM frontend achieves the best or competitive accuracy across all layouts, improving accuracy by 13.1 points on the right-angle layout (from 9.7% to 22.8%). Correlation analyses show that spatial weights align more strongly with SNR than with channel--source distance, and that higher SNR--weight correlation corresponds to higher SEC accuracy. These results demonstrate that a reconstruct-then-project, physics-based preprocessing effectively complements learning-only methods for DMAS under layout-open configurations and severe channel degradation.

声事件分类物理信息信号修复多通道传感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。