为结构化音频描述设计可验证的多维评估框架
An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

- 按标签集、描述、逻辑推理等五个维度评估结构化音频描述
- 通过可控扰动测试,准确区分语义保留与真实错误
- 适合研究音频生成与评估的学者及模型开发者
近年来,自动音频字幕(AAC)从单一句子生成转向结构化格式,显式分离声学与语义属性。然而,对这类异构数据的评估仍具挑战性。现有指标仅关注扁平文本输出,难以可靠评估多模态属性。为此,我们提出针对结构化音频描述的多轴评估框架。基于AudioCards数据集,在标签集、描述、逻辑推理、数值测量和频谱特征五个正交维度上进行评估。方法结合大语言模型(LLM)判断语义细微差别与确定性计算指标精确度量声学偏差。为严格验证框架可靠性,引入受控扰动测试协议,向真实标注中注入类型化、分级错误。结果表明,该框架能有效区分语义保持的改写与真实的语义及声学失真。
原文摘要 · Abstract (English)
Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluating this heterogeneous data remains a significant challenge. Existing caption metrics focus on flat textual outputs and fail to reliably assess multimodal attributes. To bridge this gap, we propose a multi-axis evaluation framework tailored for structured audio descriptions. Building on the AudioCards dataset, we evaluate outputs across five orthogonal axes: tag-sets, descriptions, logical reasoning, numeric measurements, and spectral profiles. Our approach combines Large Language Model (LLM) judges to capture semantic nuance with deterministic computational metrics to precisely measure acoustic deviations. To rigorously validate the reliability of this framework, we introduce a controlled perturbation testing protocol that injects typed, graded errors into groundtruth annotations. Our results demonstrate that this framework successfully distinguishes meaning-preserving paraphrases from genuine semantic and acoustic corruptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。