arXiv:2606.07229cs.SDcs.CL2026-06被引 3

首个面向通用指令的音频编辑评测基准,覆盖7类音频与多层复杂任务。

MMAE: A Massive Multitask Audio Editing Benchmark

论文配图:MMAE: A Massive Multitask Audio Editing Benchmark
图 1 · 摘自论文原文
  • 构建跨7类音频的多任务评测体系,支持从基础操作到多轮推理的6级复杂度
  • 2000个高保真样本配以17741条可验证标准,实现精准多维评估
  • 发现当前模型在混合模态任务中准确率仅0%,暴露执行精度与结构鲁棒性短板

我们提出MMAE,首个面向通用指令式音频编辑的综合性评测基准。随着智能创作兴起,交互式编辑正从视觉领域(如Nano-banana 2、Gemini-Omni)快速拓展至音频领域,但现有评估体系严重滞后,局限于特定子领域或基础操作。MMAE突破此局限,覆盖声效、语音、音乐及其混合等7种音频模态,建立涵盖6级任务复杂度(从基础修改到多跳推理)、2级粒度和8类操作类型的完整分类体系。通过人机协作精心构建的2000个高保真样本,结合首创的基于评分标准的评估框架,将自由形式任务拆解为17741条可验证准则,实现对指令遵循与上下文一致性的多维度精准评估。对主流模型的广泛测试显示,当前系统尚未达到可靠编辑水平:精确匹配率(EMR)始终低于5%,在复杂混合模态任务中更降至0%,暴露出精准执行与结构鲁棒性的关键瓶颈。MMAE旨在推动智能创作社区发展,提供清晰诊断路径,并建立标准化、可持续的下一代音频编辑评估范式。

原文摘要 · Abstract (English)

We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing. Spurred by the shift toward intelligent creation, interactive editing has rapidly expanded from visual domains, pioneered by models like Nano-banana 2 for images and Gemini-Omni for video, into audio. However, the current evaluation infrastructure lags severely, remaining highly fragmented and restricted to specific subdomains or basic operations. Unlike existing benchmarks that are limited in scope, MMAE extends to a broad spectrum of real-world scenarios, encompassing 7 distinct audio modalities, including sound, speech, music, and their mixtures. Furthermore, we establish a comprehensive taxonomy spanning 6 levels of task complexity, from basic modifications to multi-hop reasoning and multi-round editing, 2 levels of granularity, and 8 distinct operation types. Meticulously curated through human-agent collaboration, MMAE comprises 2,000 high-fidelity samples paired with a pioneering rubric-based evaluation framework. By decomposing free-form tasks into 17,741 verifiable criteria, this robust rubric-based paradigm enables a precise, multi-dimensional assessment of both instruction following and context consistency. Our extensive evaluation of leading models reveals that current systems remain far from achieving reliable edits. Strikingly, the Exact Match Rate (EMR) consistently falls below 5% and plummets to an absolute 0% in complex, mixed-modality tasks, exposing critical bottlenecks in precise execution and structural robustness. We hope MMAE will serve as a catalyst for future advances in the intelligent creation community, providing a clear diagnostic roadmap and establishing a standardized, long-lasting evaluation paradigm for next-generation audio editing systems.

音频编辑评测基准多任务指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。