arXiv:2603.07584cs.SDcs.LG2026-03被引 2

用分析驱动方法生成带精确控制标注的发动机音频数据集。

Analysis-Driven Procedural Generation of an Engine Sound Dataset with Embedded Control Annotations

  • 从真实录音中提取谐波结构,驱动参数化合成器生成新音频。
  • 每段5-10分钟音频扩展15-30倍,总时长19.0小时,含精准转速/扭矩标注。
  • 适合研究发动机音色分析、控制参数估计与神经音频生成。

计算引擎声音建模在汽车音频领域至关重要,尤其用于主动声学设计和虚拟原型开发。新兴的数据驱动引擎声音合成方法需要大量标准化、干净的音频记录,并配有精确的时间对齐运行状态标注,但这类数据因成本高、需专用测量设备且易受噪声干扰而难以获取。本文提出一种分析驱动框架,可生成带样本级控制标注的引擎音频。该方法通过自适应调谐的频谱分析从真实录音中提取谐波结构,进而驱动扩展的参数化谐波加噪声合成器。利用该框架,我们以多样化控制轨迹和参数变化,将每台发动机的原始音频(5-10分钟)扩充15-30倍,生成了过程式发动机声音数据集(Procedural Engine Sounds Dataset,19.0小时,共5,935个文件),包含跨广泛工况、信号复杂度和谐波特征的样本级转速(RPM)与扭矩标注。与真实录音对比显示,合成数据保留了典型的谐波结构;基于该数据集训练的基准可微分合成网络验证了其在数据驱动引擎声音建模中的适用性。数据集已公开,支持发动机音色分析、控制参数估计与神经生成合成研究。

原文摘要 · Abstract (English)

Computational engine sound modeling is central to the automotive audio industry, particularly for active sound design applications and virtual prototyping. Emerging data-driven engine sound synthesis methods require large volumes of standardized, clean audio recordings with precisely time-aligned operating-state annotations: data that is difficult to obtain due to high costs, specialized measurement equipment requirements, and inevitable noise contamination. We present an analysis-driven framework for generating engine audio with sample-accurate control annotations. The method extracts harmonic structures from real recordings through pitch-adaptive spectral analysis, which then drive an extended parametric harmonic-plus-noise synthesizer. With this framework, we augment 5-10 min of source audio per engine 15-30x via diverse control trajectories and parametric variation, producing the Procedural Engine Sounds Dataset (19.0 h, 5,935 files): a set of engine audio signals with sample-accurate RPM and torque annotations spanning a wide range of operating conditions, signal complexities, and harmonic profiles. Comparison against real recordings validates that the synthesized data preserves characteristic harmonic structures, and a baseline differentiable synthesis network trained on the dataset confirms its suitability for data-driven engine sound modeling. The dataset is released publicly to support research on engine timbre analysis, control parameter estimation, and neural generative synthesis.

音频生成发动机声音数据集控制标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。