测试统一音频模型生成与理解是否自洽,发现多数模型表现不佳。
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

- 设计48个三阶段自一致性测试,覆盖语音、声音、音乐共432道题。
- 最佳统一模型答对50.5%,远低于分步式基准的63.2%。
- 首次揭示统一模型在音频编辑任务中自洽性严重不足,适合评估未来音频系统。
能够实现音频理解、生成甚至编辑的统一音频模型迅速发展,但一个基础问题仍未解决:同一模型的两个分支对其生成内容是否一致?当前做法分别在专用数据集上评估各项能力,却从不检验模型能否理解自身生成结果。我们提出TORUS,首个针对原生音频统一模型的自一致性测试。TORUS包含48个三阶段测试,涵盖语音、声音、音乐,共432道六选一题目,覆盖五个任务类别。我们全面评估了五种开源统一模型,以及一个结合顶尖专用生成、编辑和理解模型的级联基线。最佳统一模型答对率为50.5%,级联基线为63.2%,随机猜测基线为16.7%。模型在音频编辑任务中表现尤为薄弱。在所有评估的模型(包括专用与统一)中,自一致性普遍有限,因此我们主张将自一致性作为未来音频系统的核心评测标准。
原文摘要 · Abstract (English)
Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound and music across five task families. We holistically evaluate five open unified models alongside a Cascaded Baseline that combines state-of-the-art specialized generation, editing and understanding models. The best unified model answers 50.5% of questions against the Cascaded Baseline's 63.2% and a 16.7% chance floor. Models struggle on audio editing. Among the evaluated audio models (specialized and unified), we observe limited self-coherence, and thus position self-coherence as an essential test for future audio systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。