用可微分框架实现物理驱动的模态声音合成与逆问题求解
DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
- 基于隐式形状表示和高阶有限元分析,构建可微分声音渲染流程
- 能准确还原目标声音,并支持物理参数、形状与撞击位置的逆向推断
- 适用于需要物理真实性声音生成的视觉、图形与机器人任务
从真实世界的声音记录中精确估计并模拟物体的物理特性,在视觉、图形与机器人领域具有重要应用价值。然而,现有进展受限:以往可微分刚体或柔体仿真方法因音频采样率过高难以直接用于模态声音合成;而传统音频合成器又未能充分建模发声物体的真实物理属性。本文提出DiffSound,一种基于隐式形状表示、新型高阶有限元分析模块及可微分音频合成器的物理驱动模态声音合成可微分渲染框架。该框架通过端到端可微性,可解决多种逆问题,包括物理参数估计、几何形状推理与撞击位置预测。实验表明,该方法能以物理合理的方式精准复现目标声音,为各类声音合成与分析应用提供有力工具。
原文摘要 · Abstract (English)
Accurately estimating and simulating the physical properties of objects from real-world sound recordings is of great practical importance in the fields of vision, graphics, and robotics. However, the progress in these directions has been limited -- prior differentiable rigid or soft body simulation techniques cannot be directly applied to modal sound synthesis due to the high sampling rate of audio, while previous audio synthesizers often do not fully model the accurate physical properties of the sounding objects. We propose DiffSound, a differentiable sound rendering framework for physics-based modal sound synthesis, which is based on an implicit shape representation, a new high-order finite element analysis module, and a differentiable audio synthesizer. Our framework can solve a wide range of inverse problems thanks to the differentiability of the entire pipeline, including physical parameter estimation, geometric shape reasoning, and impact position prediction. Experimental results demonstrate the effectiveness of our approach, highlighting its ability to accurately reproduce the target sound in a physics-based manner. DiffSound serves as a valuable tool for various sound synthesis and analysis applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。