为音效生成设计可落地的评估框架,兼顾真实感与可控性。
A Production-Oriented Framework for Evaluation of SFX Generation

- 构建参考引导的音效变体生成评估流程,统一测试标准。
- AudioX在保持音效特征与多样性上表现最优,支持音效融合。
- 适合工业音频生产团队、音效设计师使用,指导模型选型。
工业级音效设计需要既能生成逼真音频,又能保留参考音效感知身份、支持可控变化且适合实际工作流的音频生成系统。现有评估多限于文本到音频、无条件或特定任务场景,难以评估参考引导的音效变体生成。为此,本文提出面向生产的音效生成评估框架,识别出九项生产需求,明确区分模型能力差异,实现统一目标下的横向比较。框架包含两阶段协议:(1) 参考引导的音频到音频(ATA)变体任务,所有方法在相同的ESC-50音效适配设置下评估;(2) 针对原生操作的能力分析,包括音效融合、时序与能量对齐、补全和定向编辑。结合客观指标(含FAD、ImageBind参考对齐、生成变体多样性)与人类感知实验(身份保留与瞬态诊断),研究揭示了不同基线在各类生产需求下的互补优势与权衡。在共享ATA设定下,AudioX在参考对齐与多样性之间取得最佳综合平衡,同时支持音效融合;其他基线更适用于特定编辑操作。该框架为参考引导音效变体生成提供结构化评估与决策依据,并为未来一体化工业音频生成管线设计奠定基础。音频演示见附带网页。
原文摘要 · Abstract (English)
Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluations are usually tied to text-to-audio (TTA), unconditional, or task-specific settings, limiting assessment for reference-guided sound effects (SFX) variation. To address this gap, we present a production-oriented evaluation framework for structured comparison of heterogeneous audio generation and editing methods. Our framework identifies nine production requirements and explicitly accounts for differences in model capabilities, enabling comparison under a common production objective. A two-stage protocol is introduced: (1) a reference-guided audio-to-audio (ATA) variation task, in which all methods are evaluated under the same ESC-50 SFX adaptation setup, and (2) capability-specific analyses of native operations such as SFX morphing, temporal and energy alignment, inpainting, and targeted editing. This framework combines objective metrics (including FAD, ImageBind-based reference alignment, and diversity across generated variants), together with a human study of perceptual identity preservation and transient diagnosis. Our study reveals complementary strengths and trade-offs across baselines for different production needs. Among the full-generation baselines evaluated under a shared ATA setting, AudioX provides the strongest overall trade-off between reference alignment and diversity while still supporting SFX morphing. Other baselines remain most suitable for specific editing operations. Our framework establishes a structured evaluation and decision protocol for reference-guided SFX variation and provides a practical basis for designing future unified industrial audio generation pipelines. Audio demos are on the accompanying web page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。