arXiv:2603.09391cs.SDcs.AI2026-03被引 2

用可微脉冲序列建模发动机声音,更真实且参数可解释。

Physics-Informed Neural Engine Sound Modeling with Differentiable Pulse-Train Synthesis

  • 直接建模排气脉冲形状与时序,结合物理规律生成音频。
  • 相比基线模型,谐波重建提升21%,总损失降低5.7%。
  • 适合声音合成、汽车仿真与物理建模研究者使用。

发动机声音源自连续的排气压力脉冲,而非持续的谐波振荡。现有神经合成方法多聚焦于逼近最终频谱特征,本文提出直接建模底层脉冲形态与时间结构。我们提出脉冲序列共振器(Pulse-Train-Resonator, PTR)模型,一种可微合成架构,将参数化脉冲序列按发动机点火模式对齐,并通过递归Karplus-Strong共振器模拟排气声学传播。该架构融入物理先验知识,包括谐波衰减、热力学音高调制、气门动力学包络、排气系统共振,以及推导出的发动机工况(如油门操作与减速断油,DFCO)。在三种不同发动机类型共7.5小时音频上验证,PTR相较谐波加噪声基线模型实现21%的谐波重建提升和5.7%的总损失降低,同时提供对应物理现象的可解释参数。代码、模型权重与音频示例已开源。

原文摘要 · Abstract (English)

Engine sounds originate from sequential exhaust pressure pulses rather than sustained harmonic oscillations. While neural synthesis methods typically aim to approximate the resulting spectral characteristics, we propose directly modeling the underlying pulse shapes and temporal structure. We present the Pulse-Train-Resonator (PTR) model, a differentiable synthesis architecture that generates engine audio as parameterized pulse trains aligned to engine firing patterns and propagates them through recursive Karplus-Strong resonators simulating exhaust acoustics. The architecture integrates physics-informed inductive biases including harmonic decay, thermodynamic pitch modulation, valve-dynamics envelopes, exhaust system resonances and derived engine operating modes such as throttle operation and Deceleration Fuel Cutoff (DFCO). Validated on three diverse engine types totaling 7.5 hours of audio, PTR achieves a 21% improvement in harmonic reconstruction and a 5.7% reduction in total loss over a harmonic-plus-noise baseline model, while providing interpretable parameters corresponding to physical phenomena. Complete code, model weights, and audio examples are openly available.

声音合成可微分建模物理引擎发动机仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。