arXiv:2510.07978cs.AIcs.CL2025-10被引 11

评测语音助手在真实任务中的智能表现,发现其多工具协作与安全鲁棒性仍不足。

VoiceAgentBench: Are Voice Assistants ready for agentic tasks?

  • 构建6000+跨语言语音指令数据集,模拟多步骤任务和对话场景
  • 语音模型在英语任务中准确率达60.6%,印地语等语言性能显著下降
  • 揭示现有系统在复杂流程和安全测试中的普遍短板,适合评估语音智能系统

大规模语音语言模型使语音助手能理解自然口语并执行复杂任务。然而,现有语音基准多聚焦单一能力如转录或问答,未系统评估代理行为或对抗鲁棒性。为此,我们提出VoiceAgentBench,一个面向真实语音代理场景的综合性评测基准,包含6,000+合成语音指令,覆盖单工具调用、多工具工作流、多轮对话及安全性评估,涵盖英语和六种印地语系语言。为确保说话人多样性,采用基于说话人嵌入的新采样策略,优化语音转换的声学多样性。评测指标包括工具选择准确率、结构一致性及调用正确性,含对抗鲁棒性。结果显示,ASR-LLM流水线在英语任务中平均参数填充准确率达60.6%,优于端到端语音语言模型;而语音模型在印地语系语言中表现更差,且在序列化工作流与安全评估中均表现不佳,暴露出工具编排、多语言泛化与安全鲁棒性的持续缺陷。VoiceAgentBench已公开于Hugging Face:https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench,代码库发布于https://github.com/ola-krutrim/VoiceAgentBench。

原文摘要 · Abstract (English)

Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks largely focus on isolated capabilities such as transcription or question answering and do not systematically evaluate agentic behavior or adversarial robustness. To address this, we introduce VoiceAgentBench, a comprehensive benchmark for evaluating SpeechLMs in realistic spoken agentic settings, comprising 6,000+ synthetic spoken queries spanning single-tool invocations, multi-tool workflows, multi-turn dialogue, and safety evaluations across English and six Indic languages. To ensure speaker diversity, we further simulate speaker variability using a novel sampling strategy that selects audios for TTS voice conversion based on speaker embeddings to maximize acoustic diversity. Our evaluation measures tool selection accuracy, structural consistency, and the correctness of tool invocations, including adversarial robustness. Across agentic tasks, ASR-LLM pipelines outperform end-to-end SpeechLMs, achieving up to 60.6% average parameter-filling accuracy on English, while SpeechLMs exhibit lower performance and sharper degradation on Indic languages. All models struggle in sequential workflows and safety evaluations, highlighting persistent limitations in tool orchestration, multilingual generalization, and safety robustness. VoiceAgentBench is publicly available on Hugging Face at https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench, and the codebase is released at https://github.com/ola-krutrim/VoiceAgentBench.

语音智能评测基准多语言代理系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。