arXiv:2606.19595cs.LGcs.AI2026-06被引 3

评测语音助手在中断后能否正确恢复流程,发现闭源模型更稳健。

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

论文配图:IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
图 1 · 摘自论文原文
  • 构建10个企业场景的中断恢复评测基准,设计六类可控中断
  • 闭源模型任务完成率更高,长对话下性能衰减慢3.3倍
  • 适合评估客服、医疗等多步骤语音系统的鲁棒性

部署在结构化工作流(如客户服务、医疗预约、账户管理)中的语音助手需频繁处理用户打断,同时保持多步流程进度。现有语音模型评测聚焦打断时机:抢话检测、端点判断和轮次交互动态,但未衡量打断后的恢复表现:助手能否回到正确步骤?是否回应了用户插话?是否重复播报已听内容?本文提出IHBench(中断处理评测基准),评估语音助手在跨10个企业领域、状态机驱动的工作流中应对中断后的恢复能力。在对话中途注入六种类型中断,每类配备独立评分标准。每项中断从任务完成度与恢复质量两维打分。评测27个来自OpenAI、Google及开源社区的音频-语言模型配置。结果表明模型表现差异显著,恢复质量高度依赖中断类型。封闭权重模型整体更鲁棒:任务完成率更高,对话越长性能下降速度慢约3.3倍,且无音频/文本模态差距;而开源模型在三项指标上均持续退化。人工研究验证大模型判别器与人工标注者一致,跨基准分析表明恢复质量是独立的能力维度。

原文摘要 · Abstract (English)

Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user interruptions while maintaining progress through multi-step procedures. Existing benchmarks for speech-capable models focus on the timing of interruptions: barge-in detection, endpointing, and turn-taking dynamics. They leave unmeasured what happens after the interruption: does the agent resume the workflow at the correct step? Does it address the user's interjection? Does it avoid re-delivering content the user already heard? We introduce IHBench (Interruption Handling Benchmark), a benchmark that evaluates post-interruption recovery in voice agents executing state-machine-driven workflows across 10 enterprise domains. Six interruption types are injected at controlled points mid-utterance, with per-interruption evaluation rubrics generated alongside the data. Each interruption is scored on two axes: task fulfillment and recovery quality. We evaluate 27 audio-language model configurations from OpenAI, Google, and the open-weight community. Models vary widely, and recovery quality depends strongly on the interruption type. Across our experiments, closed-weight models are consistently more robust to interruptions than open-weight ones: they win far more often on task fulfillment, degrade roughly 3.3x more slowly as conversations grow longer, and show no audio-versus-text modality gap, whereas the open-weight models lose ground on all three. A human study validates the LLM judge against human annotators, and a cross-benchmark analysis against AudioMultiChallenge indicates that recovery quality is a largely distinct capability axis.

语音助手中断恢复评测基准闭源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。