构建多属性语音编辑评测基准,诊断大模型指令编辑能力瓶颈
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

- 设计七类原子与复合编辑任务,支持中英双语评测
- 提出锚定评估协议,分离计算编辑成功率与内容保留率
- 发现闭源模型更优,复合编辑仍是主要挑战
指令引导的语音编辑要求模型修改指定语音属性的同时保持非目标特征不变。尽管语音大模型(Speech LLMs)发展迅速,但该能力的系统性评估仍具挑战,因现有基准分散于孤立的编辑任务中。为此,我们提出SpeechEditBench,一个面向指令引导语音编辑的双语多属性基准。该基准包含七种原子编辑任务及整合多项操作的复合编辑任务。我们提出基于锚点的评估协议,分别衡量目标属性修改成功度与非目标语言内容保留成功度,生成三项指标:目标成功率、保留成功率和联合成功率。利用该基准,我们评估主流Speech LLMs与专用语音编辑系统。结果揭示三个关键发现:(1) 没有单一模型在所有编辑维度表现优异;(2) 闭源Speech LLMs总体优于开源模型;(3) 复合编辑构成显著挑战,即使最先进的模型也难以获得高联合成功率。SpeechEditBench为识别Speech LLMs中的瓶颈提供了严谨诊断框架,推动下一代具备更精确指令引导编辑能力的模型发展。数据与代码已公开于https://github.com/daxintan-cuhk/SpeechEditBench。
原文摘要 · Abstract (English)
Instruction-guided speech editing requires a model to modify specified speech attributes while preserving non-target characteristics. Despite rapid progress in Speech Large Language Models (Speech LLMs), systematic evaluation of this capability remains challenging, as existing benchmarks are fragmented across isolated editing tasks. To bridge this gap, we introduce SpeechEditBench, a bilingual multi-attribute benchmark for instruction-guided speech editing. SpeechEditBench encompasses seven atomic editing tasks, as well as compositional editing tasks that integrate multiple operations within a single instruction. We propose an anchor-based evaluation protocol that separately assesses the edit success of target attributes and the preservation of non-target linguistic content, leading to three metrics: target success, preservation success, and joint success. Using this benchmark, we evaluate mainstream Speech LLMs and specialized speech editing systems. The results reveal three key findings: (1) no single model performs well across all editing dimensions; (2) closed-source Speech LLMs generally outperform open-source models; (3) compositional editing poses a significant challenge, with even the most advanced models struggling to achieve high joint success. SpeechEditBench provides a rigorous diagnostic framework to identify bottlenecks in Speech LLMs, thereby facilitating the development of next-generation models with more precise instruction-guided editing capabilities. Data and code are available at https://github.com/daxintan-cuhk/SpeechEditBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。