arXiv:2508.11966cs.SD2025-08被引 5

构建首个音频编辑主观评估数据集并开发自动评测工具

Towards Automatic Evaluation and High-Quality Pseudo-Parallel Dataset Construction for Audio Editing: A Human-in-the-Loop Method

  • 引入专家评审构建6300+样本的主观评估集
  • 提出AuditEval自动评测系统,精准匹配人工评分
  • 用专家反馈筛选高质量伪平行数据,提升模型训练效果

音频编辑旨在根据文本描述操作音频内容,支持添加、删除或替换音频事件。尽管近期取得进展,但高质量基准数据集和全面评估指标的缺失仍是主要挑战。本文提出一种融合专家知识的音频编辑方法:1)建立首个主观评估数据集AuditScore,包含来自7个主流音频编辑框架和23种系统配置生成的超过6300个编辑样本,每个样本由专业评价者在整体质量、意图相关性、原特征忠实度三个维度标注;2)基于该数据集,系统提出AuditEval系列自动MOS风格评测器,覆盖基于SSL和大语言模型的方法,解决该领域缺乏有效客观指标和主观评估成本高的问题;3)利用AuditEval评估并筛选大量合成混合编辑对,通过选择最合理样本挖掘出高质量伪平行数据子集。实验验证了专家指导过滤策略显著提升数据质量,同时揭示传统客观指标的局限性和AuditEval的优势。代码与数据可访问:https://github.com/NKU-HLT/AuditEval。

原文摘要 · Abstract (English)

Audio editing aims to manipulate audio content based on textual descriptions, supporting tasks such as adding, removing, or replacing audio events. Despite recent progress, the lack of high-quality benchmark datasets and comprehensive evaluation metrics remains a major challenge for both assessing audio editing quality and improving the task itself. In this work, we propose a novel approach for audio editing task by incorporating expert knowledge into both the evaluation and dataset construction processes: 1) First, we establish AuditScore, the first comprehensive dataset for subjective evaluation of audio editing, consisting of over 6,300 edited samples generated from 7 representative audio editing frameworks and 23 system configurations. Each sample is annotated by professional raters on three key aspects of audio editing quality: overall Quality, Relevance to editing intent, and Faithfulness to original features. 2) Based on this dataset, we systematically propose AuditEval, a family of automatic MOS-style evaluators tailored for audio editing, covering both SSL-based and LLM-based approaches. It addresses the lack of effective objective metrics and the prohibitive cost of subjective evaluation in this field. 3) We further leverage AuditEval to evaluate and filter a large amount of synthetically mixed editing pairs, mining a high-quality pseudo-parallel subset by selecting the most plausible samples. Comprehensive experiments validate that our expert-informed filtering strategy effectively yields higher-quality data, while also exposing the limitations of traditional objective metrics and the advantages of AuditEval. The dataset, codes and tools can be found at: https://github.com/NKU-HLT/AuditEval.

音频编辑自动评测伪平行数据主观评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。