arXiv:2607.04337cs.SDeess.AS2026-07

构建音效真伪匹配基准,测试合成音效与真实音效的辨识能力。

Doppelganger: Sound Effects and Their Synthetic Twins

论文配图:Doppelganger: Sound Effects and Their Synthetic Twins
图 1 · 摘自论文原文
  • 构建10420对真实与合成音效配对,建立跨真实-合成边界的匹配基准。
  • 模型在未见音效上可80%准确找回原始真实录音,显著优于未训练模型。
  • 匹配能力依赖特定生成器,对不同生成器或文本驱动生成无效。

当前音频条件生成器能从真实录音生成合成音效,导致真实与合成音效在音效库和训练数据中并存,但缺乏衡量表示能否将合成片段匹配到其对应真实录音的基准。本文提出 Doppelganger 基准,包含 10,420 对真实音效(涵盖34类日常声音事件)及其对应的音频条件合成孪生体,外加一个受控的7类语料库。现成音频编码器无法干净地跨越真实-合成边界。以音效类别标签训练编码器虽在熟悉声音上有效,但在新声音上反而劣于未训练编码器。而以每对音效及其合成孪生体共同训练则表现更优:在未见音效上,模型约80%概率准确还原原始真实来源(未训练时为61%,随机水平为0.03%),且无其他指标能显著提升类别识别性能。该匹配能力具有生成器特异性——仅在相同生成器设置下有效,换用不同生成器即失效,对纯文本生成器亦不适用。人类标注基线(49人)表现高于随机但低于模型。合成孪生体有29%被误认为真实,但特定生成器检测器可完全区分合成与真实音频。

原文摘要 · Abstract (English)

Audio-conditioned generators now produce synthetic sound effects from real recordings, so the real and synthetic versions of an event increasingly coexist in sound libraries and in the corpora used to train audio models -- yet no benchmark measures whether a representation can match a synthetic clip to the specific real recording it was generated from. I introduce Doppelganger, a benchmark for matching sound effects across the synthetic-real boundary, pairing 10,420 real clips across 34 everyday sound events each with an audio-conditioned synthetic twin, alongside a controlled 7-class corpus. Off-the-shelf audio encoders do not cross the boundary cleanly. Making the embedding ignore the boundary the standard way -- training it on sound-event labels -- works on familiar sounds but backfires on new ones, dropping below the untrained encoder. Training on the pairs instead -- a clip and its own synthetic twin -- generalizes. On sound events held out of training, it recovers the exact real source about 80% of the time (up from 61% untrained; chance 0.03%), whereas no objective meaningfully improves category-level recognition on those unseen events. The learned matching is specific to one generator -- it survives changes to that generator's settings but not a switch to a different generator, and collapses for the text-only generators tested. A human annotation baseline (49 listeners) lands well above chance but below the models on the same trials. Synthetic twins fool people into calling them real about 29% of the time, yet a generator-specific detector separates these audio-conditioned twins from real recordings perfectly.

音效生成真伪识别音频基准合成音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。