用语音信号实现无需校准的室内定位,精度达1.25米。
Localization using Angle-of-Arrival Triangulation
- 基于麦克风阵列语音信号,通过角度三角测量定位说话人。
- 实测定位误差中位数为1.25米,角度估计误差仅2.2度。
- 无需用户配合或硬件改造,适合智能家居隐私保护场景。
室内定位是移动计算中的长期挑战,对智能环境(如家庭、办公室、零售空间)中位置感知与智能应用具有重要意义。随着亚马逊Alexa和谷歌Nest等语音助手日益普及,配备麦克风的智能设备正成为日常生活与家庭自动化的核心。本文提出一种被动、轻量级的基础设施定位系统,利用两个或更多空间分布的智能设备捕捉的语音信号,实现对说话人的定位。所提方法GCC+在广义互相关相位变换(GCC-PHAT)基础上,估计各设备处的到达角(AoA),并采用鲁棒三角测量技术推断说话人二维位置。为进一步提升时间分辨率与定位精度,引入特征空间扩展与子采样插值技术以精确估计到达时间差(TDoA)。系统无需硬件改造、预先校准、用户主动配合或知晓说话内容,具备高度实用性。在真实家庭环境中实验显示,角度估计中位误差为2.2度,定位中位误差为1.25米,验证了基于音频的定位在实现上下文感知、隐私保护的环境智能中的可行性与有效性。
原文摘要 · Abstract (English)
Indoor localization is a long-standing challenge in mobile computing, with significant implications for enabling location-aware and intelligent applications within smart environments such as homes, offices, and retail spaces. As AI assistants such as Amazon Alexa and Google Nest become increasingly pervasive, microphone-equipped devices are emerging as key components of everyday life and home automation. This paper introduces a passive, infrastructure-light system for localizing human speakers using speech signals captured by two or more spatially distributed smart devices. The proposed approach, GCC+, extends the Generalized Cross-Correlation with Phase Transform (GCC-PHAT) method to estimate the Angle-of-Arrival (AoA) of audio signals at each device and applies robust triangulation techniques to infer the speaker's two-dimensional position. To further improve temporal resolution and localization accuracy, feature-space expansion and subsample interpolation techniques are employed for precise Time Difference of Arrival (TDoA) estimation. The system operates without requiring hardware modifications, prior calibration, explicit user cooperation, or knowledge of the speaker's signal content, thereby offering a highly practical solution for real-world deployment. Experimental evaluation in a real-world home environment yields a median AoA estimation error of 2.2 degrees and a median localization error of 1.25 m, demonstrating the feasibility and effectiveness of audio-based localization for enabling context-aware, privacy-preserving ambient intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。