arXiv:2605.15104cs.CL2026-05

将文本工具调用评测转为语音评测,验证了模型在真实语音场景下的表现差异。

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

论文配图:From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
图 1 · 摘自论文原文
  • 用语音合成与噪声模拟生成音视频对,不重新标注即可评测语音工具调用能力。
  • 不同模型表现差异大,最高语音任务准确率达71.9,文本转语音导致性能下降1.8至4.8分。
  • 开源大模型可高精度替代人工评判,适合隐私敏感的评测场景。

语音助手越来越需要从语音中可靠调用工具,但主流评测仍基于文本。本文研究是否可在不重标注工具定义和标准答案的前提下,将验证过的文本评测转化为受控的音频工具调用评估。提出的数据集无关框架通过文本转语音、说话人差异和环境噪声,生成配对的文本-音频实例,同时保留原始标注。基于对7个全模态模型在Confetti和When2Call音频转换版本上的广泛评估,结果表明性能高度依赖模型与任务:Gemini-3.1-Flash-Live在Confetti上得分最高(70.4),GPT-Realtime-1.5在When2Call上表现最佳(71.9)。在Confetti上,文本转语音性能差距在1.8分(Qwen3-Omni)至4.8分(GPT-Realtime-1.5)之间。失败案例分析显示,性能下降主要源于对语音中参数值的理解错误。针对实际部署,进一步报告纯文本结果、基于歧义的重构压力测试,以及经人类偏好验证的无参考式大模型评分协议。值得注意的是,参数量至少80亿的开源Qwen3模型与专有模型评判者一致性超过80%,支持隐私保护型评估。总体而言,本框架提供了可验证、可复现的第一阶段诊断,补充专用音频语料库。

原文摘要 · Abstract (English)

Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be converted into controlled audio-based tool calling evaluations without re-annotating the tool schema and gold labels. Our dataset-agnostic framework uses text-to-speech, speaker variation, and environmental noise to create paired text-audio instances while preserving the original dataset annotations. Based on extensive evaluation of 7 omni-modal models on audio-converted versions of Confetti and When2Call, our framework demonstrates that the performance is strongly model- and task-dependent: Gemini-3.1-Flash-Live obtains the highest Confetti score (70.4), whereas GPT-Realtime-1.5 performs best on When2Call (71.9). On Confetti, the text-to-voice gap ranges from 1.8 points for Qwen3-Omni to 4.8 points for GPT-Realtime-1.5. A targeted analysis of failure cases demonstrates that degradations most often reflect misunderstandings of argument values in the speech. Considering real-world deployment scenarios, we further report text-only results, an ambiguity-based reformulation stress test, and a reference-free LLM-as-judge protocol validated against human preferences. Notably, we find that open-source Qwen3 judges with at least 8B parameters exceed 80% agreement with proprietary judges, supporting privacy-preserving evaluation. Overall, our framework provides a verifiable and reproducible first-stage diagnostic that complements purpose-built audio corpora.

语音评测工具调用大模型评估隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。