用事件相机捕捉物体振动,实现非接触式声音还原。
EvMic: Event-based Non-contact sound recovery from effective spatial-temporal modeling
- 构建事件流时空联合建模框架,融合稀疏事件与长时序特征。
- 在合成与真实数据上实现高保真声波还原,提升信号质量。
- 适合对低功耗、高频率视觉声学系统感兴趣的研究者。
当声波撞击物体时,会引发高频微小的振动,进而导致视觉上的细微变化,可用于声音恢复。早期研究常受限于采样率、带宽、视场范围以及光学路径复杂性之间的权衡。近年来,事件相机硬件的进步展现出在视觉声学恢复中的潜力,因其对高频信号的捕捉能力优异。然而,现有基于事件相机的振动恢复方法在声音还原方面仍不理想。本文提出一种全新的非接触式声音恢复流程,充分挖掘事件流中的时空信息。首先,通过新颖的仿真管道生成大规模训练数据集;其次,设计网络利用事件稀疏性捕捉空间信息,并采用Mamba模型建模长期时序依赖;最后,引入空间聚合模块整合多位置信息以进一步提升信号质量。为有效捕获声波引起的事件信号,还设计了一种基于激光阵列的成像系统以增强梯度对比度,并采集多组数据序列用于测试。在合成与真实数据上的实验结果验证了该方法的有效性。
原文摘要 · Abstract (English)
When sound waves hit an object, they induce vibrations that produce high-frequency and subtle visual changes, which can be used for recovering the sound. Early studies always encounter trade-offs related to sampling rate, bandwidth, field of view, and the simplicity of the optical path. Recent advances in event camera hardware show good potential for its application in visual sound recovery, because of its superior ability in capturing high-frequency signals. However, existing event-based vibration recovery methods are still sub-optimal for sound recovery. In this work, we propose a novel pipeline for non-contact sound recovery, fully utilizing spatial-temporal information from the event stream. We first generate a large training set using a novel simulation pipeline. Then we designed a network that leverages the sparsity of events to capture spatial information and uses Mamba to model long-term temporal information. Lastly, we train a spatial aggregation block to aggregate information from different locations to further improve signal quality. To capture event signals caused by sound waves, we also designed an imaging system using a laser matrix to enhance the gradient and collected multiple data sequences for testing. Experimental results on synthetic and real-world data demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。