arXiv:2511.19396cs.SDcs.AI2025-11

用手机级芯片实现实时声源追踪,动态调整麦克风方向。

Real-Time Object Tracking with On-Device Deep Learning for Adaptive Beamforming in Dynamic Acoustic Environments

  • 用单目深度+双目视觉实现3D目标定位
  • 声学阵列实时跟随目标,信干比显著提升
  • 适合智能会议、智能家居等实时场景

对象跟踪与声学波束成形技术的进步正推动监控、人机交互和机器人领域的创新。本文提出一种嵌入式系统,将基于深度学习的跟踪与波束成形结合,在动态声学环境中实现精准声源定位与定向音频采集。该方法利用单摄像头深度估计与双目视觉,实现移动物体的准确三维定位。采用由MEMS麦克风构成的平面同心圆麦克风阵列,提供紧凑、低功耗的2D波束转向平台,支持方位角与俯仰角调节。实时跟踪输出持续调整阵列聚焦方向,使声学响应与目标位置同步。通过融合学习所得的空间感知与动态转向能力,系统在多源或移动源环境下仍保持稳健性能。实验评估显示信干比显著提升,表明该设计非常适合远程会议、智能家庭设备及辅助技术应用。

原文摘要 · Abstract (English)

Advances in object tracking and acoustic beamforming are driving new capabilities in surveillance, human-computer interaction, and robotics. This work presents an embedded system that integrates deep learning-based tracking with beamforming to achieve precise sound source localization and directional audio capture in dynamic environments. The approach combines single-camera depth estimation and stereo vision to enable accurate 3D localization of moving objects. A planar concentric circular microphone array constructed with MEMS microphones provides a compact, energy-efficient platform supporting 2D beam steering across azimuth and elevation. Real-time tracking outputs continuously adapt the array's focus, synchronizing the acoustic response with the target's position. By uniting learned spatial awareness with dynamic steering, the system maintains robust performance in the presence of multiple or moving sources. Experimental evaluation demonstrates significant gains in signal-to-interference ratio, making the design well-suited for teleconferencing, smart home devices, and assistive technologies.

声源定位边缘计算实时跟踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。