arXiv:2601.13742cs.CL2026-01Conference of the …被引 4

让大模型通过音频线索进行语音评估,更便宜且更贴近人类判断。

Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation

  • 将音频信号转为文本描述,让大模型基于线索推理评分
  • 在内容、音质、副语言三维度上与真人评分高度一致
  • 适合需要低成本高可解释性的语音质量评估场景

大语言模型(LLM)在文本评估中表现强劲,但无法直接处理音频。当前语音到语音(S2S)评估依赖昂贵且不透明的音频语言模型(ALMs)。本文提出TRACE框架,使LLM能够基于音频线索进行推理,实现低成本、对齐人类判断的S2S评估。我们引入人类思维链(HCoT)标注协议,将评估拆分为内容(C)、语音质量(VQ)和副语言(P)三个显性维度。基于该数据,TRACE构建廉价音频信号的文本蓝图,引导LLM逐维打分,并通过确定性策略融合为总分。实验表明,TRACE在与人类评分的一致性上优于ALMs和仅用转录文本的LLM,同时成本显著降低。相关HCoT标注与框架将开源,推动可扩展、对齐人类的语音评估发展。

原文摘要 · Abstract (English)

Large Language Model (LLM) judges exhibit strong reasoning capabilities but are limited to textual content. This leaves current automatic Speech-to-Speech (S2S) evaluation methods reliant on opaque and expensive Audio Language Models (ALMs). In this work, we propose TRACE (Textual Reasoning over Audio Cues for Evaluation), a novel framework that enables LLM judges to reason over audio cues to achieve cost-efficient and human-aligned S2S evaluation. To demonstrate the strength of the framework, we first introduce a Human Chain-of-Thought (HCoT) annotation protocol to improve the diagnostic capability of existing judge benchmarks by separating evaluation into explicit dimensions: content (C), voice quality (VQ), and paralinguistics (P). Using this data, TRACE constructs a textual blueprint of inexpensive audio signals and prompts an LLM to render dimension-wise judgments, fusing them into an overall rating via a deterministic policy. TRACE achieves higher agreement with human raters than ALMs and transcript-only LLM judges while being significantly more cost-effective. We will release the HCoT annotations and the TRACE framework to enable scalable and human-aligned S2S evaluation.

语音评估大模型推理低成本人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。