首个面向可穿戴设备语音助手的真实场景评测基准,聚焦运动噪声与多说话人干扰。
WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
- 构建3842段多通道第一视角音频数据集,覆盖五类真实交互任务。
- 主流语音大模型在户外嘈杂环境准确率仅29%~59%,性能显著下降。
- 多通道输入能有效提升抗噪能力,更好区分设备指令与背景对话。
AI眼镜等可穿戴设备正将语音助手变为随时可用、无需双手的日常伙伴,但带来第一视角音频受运动与噪声影响、快速微交互以及区分设备指令与背景对话等挑战。现有评测基准大多忽略这些复杂性,集中于纯净或通用对话音频。为此,我们提出WearVox,首个专为真实可穿戴场景设计的语音助手评测基准。WearVox包含3,842段通过AI眼镜在五类任务(搜索引导问答、闭卷问答、旁白拒绝、工具调用、语音翻译)中采集的多通道第一视角音频,涵盖多种室内外环境与声学条件。每段音频均配有丰富元数据,支持对模型在真实约束下的性能进行细致分析。我们评测了领先的专有与开源语音大语言模型(SLLMs),发现多数实时SLLMs在WearVox上的准确率介于29%至59%之间,尤其在嘈杂户外环境中表现大幅下降,凸显该基准的难度与真实性。此外,我们对两个新SLLM进行了案例研究,分别使用单通道与多通道音频推理,结果表明多通道输入显著提升模型对环境噪声的鲁棒性,并增强对设备指令与背景语音的区分能力。研究强调空间音频线索对上下文感知语音助手的关键作用,并确立WearVox作为推动可穿戴语音AI研究的综合性测试平台。
原文摘要 · Abstract (English)
Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。