构建语音特效数据集,实现精准特效识别与评估。
VoxEffects: A Speech-Oriented Audio Effects Dataset and Benchmark

- 提供多粒度特效链标注的语音数据集
- 支持特效存在检测、预设分类和强度预测
- 适用于语音处理鲁棒性与公平性研究
真实场景中的语音常经过后期音效处理,但现有语音数据集缺乏精确的音效及其参数标注,限制了系统性研究。我们提出VoxEffects,一个将合成语音与多粒度特效链精确标注配对的数据集。该数据集支持语音导向的音效识别任务:给定一段合成波形,推断所应用的音效及其参数。数据集基于最小编辑的干净语音构建,提供离线合成与实时渲染的可扩展流水线,适用于高效训练与评估。音效识别基准涵盖音效存在检测、预设分类与强度预测,并包含覆盖采集端与平台端退化的鲁棒性测试协议。我们提供了基于AudioMAE的多任务基线模型,并分析了领域偏移、鲁棒性、输入时长与性别公平性问题。
原文摘要 · Abstract (English)
Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects dataset that pairs produced speech with exact effect-chain supervision at multiple granularities. VoxEffects supports speech-oriented audio effect identification: given a produced waveform, infer which effects are present and how they are applied. Built from minimally edited clean speech, it provides an extensible rendering pipeline for both offline synthesis and on-the-fly rendering for efficient training and evaluation. The audio effect identification benchmark includes effect presence detection, preset classification, and intensity prediction, with a robustness protocol covering capture-side and platform-side degradations. We provide an AudioMAE-based multi-task baseline and analyses of domain shift, robustness, input duration, and gender fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。