构建3万条语音指令数据集,评估语音助手在真实场景下的工具调用能力。
Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
- 设计跨智能家居、车载、可穿戴三领域的语音指令数据集
- 复杂指令下现有模型性能下降超40%,暴露推理短板
- 含真实噪声与语音克隆,适合测试语音助手鲁棒性
语音助手依赖语音语言模型(SpeechLM)理解口语指令并执行复杂任务,但现有评测基准缺乏领域广度、声学多样性及组合推理复杂性,难以评估工具调用表现。本文提出Audio2Tool,一个包含约3万条查询的大规模数据集,用于评估SpeechLM在智能汽车、智能家居和可穿戴设备三大领域中的工具调用能力。该基准设有多层次复杂度层级,从简单直接指令到多意图、大海捞针式信息提取,以识别不同故障模式。为保证真实性,采用零样本语音克隆的文本转语音合成及多样化噪声环境模拟真实使用场景。对前沿SpeechLM与ASR-LLM流水线的评估显示,其在简单指令上表现良好,但在组合性与声学挑战下性能显著下降。代码与数据集已公开于项目主页:https://audio2tool.github.io/。
原文摘要 · Abstract (English)
Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning complexity to evaluate tool-calling performance. We introduce Audio2Tool, a large-scale dataset comprising approximately 30,000 queries designed to assess tool-calling capabilities of SpeechLMs across three primary domains: Smart Car, Smart Home, and Wearables. Our benchmark features a multi-tier complexity hierarchy, ranging from simple direct commands to complex multi-intent and needle-in-a-haystack extraction to isolate distinct failure modes. To ensure realism, we employ zero-shot voice cloning text-to-speech synthesis and diverse noise profiles to simulate in-the-wild conditions. Evaluations of state-of-the-art SpeechLMs and ASR-LLM pipelines show strong performance on simple commands but significant degradation under compositional and acoustic challenges. Code and dataset are publicly available on the project page: https://audio2tool.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。