arXiv:2509.16975cs.SDeess.AS2025-09被引 1

用大模型生成可解释的音频编辑评估,更接近真人判断。

Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs

  • 基于Qwen2-Audio设计思维链提示,分步推理音频差异与共性。
  • 自动评分与人类判断高度一致,优于现有基线方法。
  • 适合需要透明评估结果的音频生成与修复研究者。

自动均值意见得分(MOS)预测为音频评估提供了可扩展且一致的替代方案,既避免了主观听感测试,又超越了传统客观指标。受“大模型作为评判者”范式启发,近期的多模态大语言模型展现出强大的感知建模与推理能力,可用于音频质量评估。本文针对音频编辑评估这一挑战性问题,提出首个基于自然语言的自动化评估框架,构建于Qwen2-Audio之上。通过引入两种基于字幕的微调任务以增强多音频理解能力,并设计链式思维提示策略,促使模型进行结构化、分步骤的推理。实验表明,该框架能生成可解释且逻辑一致的文本评估,与人类判断高度吻合,性能优于现有基线。代码与演示已开源:https://github.com/NKU-HLT/Eval_Reasoning。

原文摘要 · Abstract (English)

Automatic mean opinion score (MOS) prediction serves as a principled alternative to both subjective listening tests and objective metrics, providing scalable and consistent audio evaluation. Inspired by the LLM-as-Judge paradigm, recent multimodal large language models offer strong perceptual modeling and reasoning capabilities, enabling audio quality assessment. In this work, we address the challenging problem of audio editing evaluation and propose the first natural language-based automated evaluation framework built upon Qwen2-Audio. Two caption-based fine-tuning tasks are introduced to enhance multi-audio understanding, together with a designed Chain-of-Thought prompting strategy to encourage structured, step-by-step reasoning. Experiments show that our framework produces interpretable and logically consistent text-based evaluations, aligning closely with human judgments while outperforming existing baselines. The code and demo are available at https://github.com/NKU-HLT/Eval_Reasoning.

音频评估大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。