arXiv:2605.04505eess.AScs.AI2026-05被引 1

用自然语言指令让大模型零样本评估音频,效果超越现有方法。

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

论文配图:JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
图 1 · 摘自论文原文
  • 将音频评估转为自指导推理任务,结合冻结音视频编码器与微调大模型。
  • 在语音、声音、音乐等任务上相关性达最新水平,无需额外训练。
  • 适合需要灵活、通用音频评价的科研与工业场景。

生成式音频模型的快速发展已超过评估方法的进步。现有客观指标和通用多模态大语言模型(MLLMs)常面临领域泛化差、零样本能力弱及指令灵活性不足的问题。为此,我们提出JASTIN,一种可泛化的、指令驱动的音频评估框架,将音频评估建模为自指导推理任务。JASTIN通过可训练的音频适配器连接冻结的高性能音频编码器与微调的大语言模型骨干网络。为确保稳健的零样本泛化能力,我们设计了包含多源、多任务、多校准和多描述的数据构建流程。实验表明,JASTIN在与人工主观评分的皮尔逊和斯皮尔曼相关性上达到当前最优水平,且在语音、声音、音乐及跨域评估任务中均持续优于通用MLLMs,无需任务特定微调。

原文摘要 · Abstract (English)

The rapid advancement of generative audio models has outpaced the development of robust evaluation methodologies. Existing objective metrics and general multimodal large language models (MLLMs) often struggle with domain generalization, zero-shot capabilities, and instructional flexibility. To address these bottlenecks, we propose JASTIN, a generalizable, instruction-driven audio evaluation framework that formulates audio assessment as a self-instructed reasoning task. JASTIN bridges a frozen high-performance audio encoder with a fine-tuned LLM backbone via a trainable audio adapter. To ensure robust zero-shot generalization, we introduce a comprehensive instruction following data preparation pipeline, incorporating Multi-Source, Multi-Task, Multi-Calibration, and Multi-Description data. Experimental results demonstrate that JASTIN achieves state-of-the-art Pearson and Spearman correlations with human subjective ratings. It consistently outperforms general MLLMs across speech, sound, music, and out-of-domain evaluation tasks without the need for task-specific retraining.

音频评估大模型零样本指令学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。