评测智能眼镜在多人对话中的语音识别与理解能力,发现重叠说话严重影响识别效果。
The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

- 基于真实场景构建四通道106小时多说话人数据集,支持双轨评估
- 重叠说话导致语音识别准确率显著下降,语调理解仍难达标
- 适合研究可穿戴设备语音交互与多说话人处理的开发者和学者
大型语言模型(LLMs)和多模态大语言模型(MLLMs)的发展为可穿戴语音接口带来了新机遇,智能眼镜作为以佩戴者为中心的持续音频感知平台,具备独特优势。然而,在动态声学环境、说话人重叠及佩戴者中心录音几何带来的空间模糊性下,语音识别与理解仍具挑战。为此,我们推出了IEEE SLT 2026智能眼镜挑战赛,聚焦以佩戴者为中心的多说话人语音处理。挑战包含两个赛道:两人对话理解与多人会议理解,联合评估时间标记说话人归属自动语音识别(TSA-ASR)与口语理解(SLU)。数据集基于106小时四通道真实场景语音,涵盖714个会话。本文介绍任务设计、数据构建、参赛结果,并总结共享评估的主要发现。结果显示,严重说话人重叠仍是影响TSA-ASR性能的关键因素,当前音频-语言模型在复杂SLU设置中仍难以理解语调等副语言特征。更多细节见官方挑战网站。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentric platform for continuous audio sensing and assistance. However, speech recognition and understanding in this setting remain challenging because of dynamic acoustic conditions, speaker overlap, and the spatial ambiguity introduced by wearer-centered recording geometry. To support systematic evaluation in this setting, we introduce the IEEE SLT 2026 SmartGlasses Challenge for egocentric multi-speaker speech processing. The challenge consists of two tracks, Dyadic Dialogue Understanding and Multi-party Meeting Understanding, and jointly evaluates Time-Stamped Speaker-Attributed Automatic Speech Recognition (TSA-ASR) and Spoken Language Understanding (SLU). It is built on a 106-hour four-channel egocentric speech dataset containing 714 sessions collected in real-world scenarios. This paper describes challenge tasks, dataset construction, submissions, and summarizes the main findings from the shared evaluation. The results show that heavy speaker overlap remains a major factor affecting TSA-ASR performance, while paralinguistic acoustic understanding continues to be difficult for current audio-language models in complex SLU settings. Further details can be found on the official challenge website.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。