评测语音交互系统推理能力,发现语音模型性能远低于文本模型。
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
- 构建语音原生推理评测集VERA,涵盖五类任务
- 语音模型平均准确率仅11.3%,远低于文本的54.0%
- 揭示实时语音系统存在性能天花板,适合语音助手研发者参考
我们提出语音推理能力评测基准VERA,用于在实时对话约束下评估语音交互系统的推理能力。VERA包含2,931个源自经典文本基准的语音原生任务,分为数学、网络、科学、长上下文和事实五类,每项任务均保留原始推理难度并适配语音交互。VERA支持模型家族内文本与语音的直接对比,并可分析架构设计对可靠性的影响。我们评估了12个主流语音系统及强文本基线,发现显著且一致的模态差距:在竞赛数学任务中,领先文本模型准确率达74.8%,其语音版本仅为6.1%;跨赛道宏平均下,最优文本模型为54.0%,语音模型仅11.3%。延迟-精度分析显示,低延迟系统准确率稳定在约10%;接近文本性能需牺牲实时性。诊断实验表明,延长思考时间收益极微,解耦推理与叙述的级联结构虽提升准确率但仍远低于文本,并引入特定的语义一致性和定位错误。故障分析揭示了流式、端到端和级联设计的差异性错误模式。VERA提供可复现的测试平台与针对性诊断工具,为解耦思考与表达的架构提供量化的进步路径,助力实现既流畅又可靠的真实语音助手。
原文摘要 · Abstract (English)
We present Voice Evaluation of Reasoning Ability (VERA), a benchmark for evaluating reasoning ability in voice-interactive systems under real-time conversational constraints. VERA comprises 2,931 voice-native episodes derived from established text benchmarks and organized into five tracks (Math, Web, Science, Long-Context, Factual). Each item is adapted for speech interaction while preserving reasoning difficulty. VERA enables direct text-voice comparison within model families and supports analysis of how architectural choices affect reliability. We assess 12 contemporary voice systems alongside strong text baselines and observe large, consistent modality gaps: on competition mathematics a leading text model attains 74.8% accuracy while its voice counterpart reaches 6.1%; macro-averaged across tracks the best text models achieve 54.0% versus 11.3% for voice. Latency-accuracy analyses reveal a low-latency plateau, where fast voice systems cluster around ~10% accuracy, while approaching text performance requires sacrificing real-time interaction. Diagnostic experiments indicate that common mitigations are insufficient. Increasing "thinking time" yields negligible gains; a decoupled cascade that separates reasoning from narration improves accuracy but still falls well short of text and introduces characteristic grounding/consistency errors. Failure analyses further show distinct error signatures across native streaming, end-to-end, and cascade designs. VERA provides a reproducible testbed and targeted diagnostics for architectures that decouple thinking from speaking, offering a principled way to measure progress toward real-time voice assistants that are both fluent and reliably reasoned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。