arXiv:2504.08907cs.SDcs.CL2025-04被引 6

让可穿戴设备听懂声音方向,实现更智能的语音交互。

Spatial Audio Processing with Large Language Model on Wearable Devices

  • 用单麦克风通过微结构传感提取声音方向,融合语言特征。
  • 方向识别误差仅25.72°,比现有方法提升显著,词错误率5.3%。
  • 轻量化适配适合设备端运行,适合AR、无障碍等场景。

将空间上下文融入大语言模型(LLMs)有望革新人机交互,尤其在可穿戴设备上。本文提出一种新系统架构,将空间语音理解能力集成至LLMs,实现对可穿戴技术的上下文感知与自适应应用。方法基于微结构空间传感,利用单声道麦克风提取精确的方向到达(DoA)信息;为弥补缺乏相关数据集的问题,我们以LibriSpeech为基础合成名为OmniTalk的数据集。该空间信息与OpenAI Whisper模型的语言嵌入融合,使各模态学习互补的上下文表征。融合嵌入经对齐后输入LLaMA-3.2 3B模型,并采用轻量级适配技术LoRA进行微调,优化设备端处理性能。系统SING支持空间感知语音识别(ASR),平均误差达25.72°,显著优于现有工作88.52°的中位误差,词错误率(WER)为5.3%。同时支持声景分析,如推断说话人数及方向,最多支持5人,中位DoA误差16°。系统在空间语音理解方面表现优异,同时兼顾功耗效率、隐私与硬件限制,为增强现实、无障碍辅助和沉浸式体验提供技术支持。

原文摘要 · Abstract (English)

Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial speech understanding into LLMs, enabling contextually aware and adaptive applications for wearable technologies. Our approach leverages microstructure-based spatial sensing to extract precise Direction of Arrival (DoA) information using a monaural microphone. To address the lack of existing dataset for microstructure-assisted speech recordings, we synthetically create a dataset called OmniTalk by using the LibriSpeech dataset. This spatial information is fused with linguistic embeddings from OpenAI's Whisper model, allowing each modality to learn complementary contextual representations. The fused embeddings are aligned with the input space of LLaMA-3.2 3B model and fine-tuned with lightweight adaptation technique LoRA to optimize for on-device processing. SING supports spatially-aware automatic speech recognition (ASR), achieving a mean error of $25.72^\circ$-a substantial improvement compared to the 88.52$^\circ$ median error in existing work-with a word error rate (WER) of 5.3. SING also supports soundscaping, for example, inference how many people were talking and their directions, with up to 5 people and a median DoA error of 16$^\circ$. Our system demonstrates superior performance in spatial speech understanding while addressing the challenges of power efficiency, privacy, and hardware constraints, paving the way for advanced applications in augmented reality, accessibility, and immersive experiences.

语音识别可穿戴空间感知LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。