arXiv:2604.15037cs.AIcs.CL2026-04被引 1

首个评估语音助手主动性的基准,揭示大模型在主动干预上的短板

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench

论文配图:From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
图 1 · 摘自论文原文
  • 构建四类新任务,聚焦语音助手的主动行为评估
  • 测试1182个高质量样本,发现当前模型易误触发且推理弱
  • 适合研究主动交互、语音智能系统的开发者与学者

大型语言模型代理正从被动文本响应转向主动多模态交互,但现有评测体系仍以被动应答为主,忽视主动干预与持续监控的复杂性。为此,我们提出首个专为主动语音代理设计的评估框架ProVoice-Bench,包含四项全新任务。通过多阶段数据合成流程,构建了1182个高质量测试样本。对前沿多模态大模型的评估显示,其在主动触发过频和推理能力方面存在显著差距。结果凸显当前模型局限,为打造更自然、情境感知的主动代理提供了发展路径。

原文摘要 · Abstract (English)

Recent advancements in LLM agents are gradually shifting from reactive, text-based paradigms toward proactive, multimodal interaction. However, existing benchmarks primarily focus on reactive responses, overlooking the complexities of proactive intervention and monitoring. To bridge this gap, we introduce ProVoice-Bench, the first evaluation framework specifically designed for proactive voice agents, featuring four novel tasks. By leveraging a multi-stage data synthesis pipeline, we curate 1,182 high-quality samples for rigorous testing. Our evaluation of state-of-the-art Multimodal LLMs reveals a significant performance gap, particularly regarding over-triggering and reasoning capabilities. These findings highlight the limitations of current models and offer a roadmap for developing more natural, context-aware proactive agents.

语音代理主动性评估多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。