用统计方法生成和评估纹理音效,更贴近人耳感知。
A Statistics-Driven Differentiable Approach for Sound Texture Synthesis and Analysis
- 基于统计特征设计新损失函数TexStat,不依赖时间结构。
- 在多种纹理音效上验证,该方法对噪声鲁棒且感知有效。
- 适合音频生成与质量评估,代码开源可配置。
本文提出TexStat,一种专为具有随机结构和感知平稳性的纹理音效设计的新型损失函数。受McDermott与Simoncelli的统计与感知框架启发,TexStat能在不依赖时序结构的情况下识别同类别信号间的相似性。同时,我们提出将TexStat与弗雷谢音频距离(FAD)结合,作为纹理音效合成模型的评估指标。此外,我们还提出了TexEnv——一个高效、轻量且可微分的纹理音效合成器,通过在滤波噪声上施加包络生成音频。进一步地,我们将这些组件整合为面向纹理音效的生成模型TexDSP,其灵感来自DDSP。在多种纹理音效上的大量实验表明,TexStat具有感知意义、时不变性且对噪声鲁棒,使其在生成任务中作为损失函数及作为评估指标均表现优异。所有工具与代码均以开源形式提供,我们的PyTorch实现具备高效性、可微性和高度可配置性,适用于生成任务与感知基准评估。
原文摘要 · Abstract (English)
In this work, we introduce TexStat, a novel loss function specifically designed for the analysis and synthesis of texture sounds characterized by stochastic structure and perceptual stationarity. Drawing inspiration from the statistical and perceptual framework of McDermott and Simoncelli, TexStat identifies similarities between signals belonging to the same texture category without relying on temporal structure. We also propose using TexStat as a validation metric alongside Frechet Audio Distances (FAD) to evaluate texture sound synthesis models. In addition to TexStat, we present TexEnv, an efficient, lightweight and differentiable texture sound synthesizer that generates audio by imposing amplitude envelopes on filtered noise. We further integrate these components into TexDSP, a DDSP-inspired generative model tailored for texture sounds. Through extensive experiments across various texture sound types, we demonstrate that TexStat is perceptually meaningful, time-invariant, and robust to noise, features that make it effective both as a loss function for generative tasks and as a validation metric. All tools and code are provided as open-source contributions and our PyTorch implementations are efficient, differentiable, and highly configurable, enabling its use in both generative tasks and as a perceptually grounded evaluation metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。