arXiv:2601.07367cs.SD2026-01被引 2

提出新基准框架FOCAL,评估多模态语音代理的端到端推理与错误传播。

FOCAL: A Novel Benchmarking Technique for Multi-modal Agents

  • 构建端到端测试框架,支持自动化与人工辅助评测
  • 引入推理得分与语义得分,量化语音对话有效性
  • 可分析组件级错误传播,适用于语音+文本双模态代理

随着推理能力的进步,结合MCP服务器和音频语言模型(ALMs)的多模态代理(支持语音与文本)已成为行业前沿。尽管级联式语音代理因大语言模型提供的强大推理能力仍占主导地位,但其常存在错误在管道中逐级传播的问题。本文提出FOCAL框架,用于对多模态代理(语音到语音+文本输入)进行端到端推理、组件级错误传播及错误分析的基准测试。我们还提出了两项新指标:推理得分与语义得分,以评估代理在语音模式下进行有意义对话的有效性。

原文摘要 · Abstract (English)

With the recent advancements in reasoning capabilities, tool calling using MCP servers and Audio Language Models (ALMs), development and integration of multi-modal agents (with voice and text support) has come to the industry forefront. Cascading pipelines for voice agents still play a central role in the industry owing to their superior reasoning capabilities facilitated by LLMs. Although, cascading pipelines often present error propagation through the pipeline. We propose a framework, FOCAL to benchmark end-to-end reasoning, component-wise error propagation and error analysis for automated as well as human-assisted testing of multi-modal agents (voice to voice + text input). We also share two novel metrics viz. Reasoning and Semantic scores to evaluate efficacy of the agent in having meaningful conversations in voice mode.

多模态代理语音评测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。