arXiv:2603.00395cs.SDcs.LG2026-03

让助听设备像调音台一样精细调控环境声音。

Fine-grained Soundscape Control for Augmented Hearing

  • 动态界面只显示当前活跃的声音类别,用户可独立调节每类音量。
  • 支持最多5个重叠声源的实时分离,对设备算力要求低且响应快。
  • 适合需要在复杂环境中精准听声的人群,如听力障碍者或专注工作者。

助听设备日益普及,但其声音控制仍较粗略:用户只能开启全局降噪或聚焦单一声源。然而真实场景中存在多个同时发声的声源,用户可能希望分别调整。本文提出Aurchestra,首个在资源受限助听设备上实现细粒度、实时声音场景控制的系统。系统包含两个核心组件:(1) 动态界面仅呈现当前活跃的声音类别;(2) 实时、本地部署的多输出提取网络,为每个选中类别生成独立音轨,可处理最多5个重叠目标声源,在未见的室内外真实环境中实现显著的目标声增强与干扰抑制。该系统支持在6毫秒音频块上实时运行,优化适配多种计算受限平台,使用户如同调音师般自由混音环境声音。结果表明,世界不必被听成单一、无差别的流:借助Aurchestra,声音场景真正变得可编程。

原文摘要 · Abstract (English)

Hearables are becoming ubiquitous, yet their sound controls remain blunt: users can either enable global noise suppression or focus on a single target sound. Real-world acoustic scenes, however, contain many simultaneous sources that users may want to adjust independently. We introduce Aurchestra, the first system to provide fine-grained, real-time soundscape control on resource-constrained hearables. Our system has two key components: (1) a dynamic interface that surfaces only active sound classes and (2) a real-time, on-device multi-output extraction network that generates separate streams for each selected class, achieving robust performance for upto 5 overlapping target sounds, and letting users mix their environment by customizing per-class volumes, much like an audio engineer mixes tracks. We optimize the model architecture for multiple compute-limited platforms and demonstrate real-time performance on 6 ms streaming audio chunks. Across real-world environments in previously unseen indoor and outdoor scenarios, our system enables expressive per-class sound control and achieves substantial improvements in target-class enhancement and interference suppression. Our results show that the world need not be heard as a single, undifferentiated stream: with Aurchestra, the soundscape becomes truly programmable.

助听设备声音分离实时控制可编程声景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。