arXiv:2604.02390cs.SDcs.AI2026-04中稿 · publication by the…

用空间感知融合提升音视频导航效率,支持未听过的语音目标

Spatial-Aware Conditioned Fusion for Audio-Visual Navigation

  • 将目标相对方位距离离散化为紧凑描述符,用于策略条件与状态建模
  • 通过音频嵌入与空间描述符动态调制视觉特征,生成指向目标的融合表征
  • 显著提升导航效率和泛化能力,计算开销低,适合未知声音场景

音视频导航任务要求智能体仅凭视觉观测和声学线索定位并导航至持续发声的目标。现有方法多依赖简单特征拼接或晚期融合,缺乏目标相对位置的显式离散表示,限制了学习效率与泛化能力。本文提出空间感知条件融合(SACF)。SACF首先从音视频线索中离散化目标相对方向与距离,预测其分布,并编码为紧凑描述符以用于策略条件与状态建模。随后,利用音频嵌入与空间描述符生成通道级缩放与偏置,通过条件线性变换调制视觉特征,生成面向目标的融合表征。SACF在降低计算开销的同时提升了导航效率,并能良好泛化至未见过的目标声音。

原文摘要 · Abstract (English)

Audio-visual navigation tasks require agents to locate and navigate toward continuously vocalizing targets using only visual observations and acoustic cues. However, existing methods mainly rely on simple feature concatenation or late fusion, and lack an explicit discrete representation of the target's relative position, which limits learning efficiency and generalization. We propose Spatial-Aware Conditioned Fusion (SACF). SACF first discretizes the target's relative direction and distance from audio-visual cues, predicts their distributions, and encodes them as a compact descriptor for policy conditioning and state modeling. Then, SACF uses audio embeddings and spatial descriptors to generate channel-wise scaling and bias to modulate visual features via conditional linear transformation, producing target-oriented fused representations. SACF improves navigation efficiency with lower computational overhead and generalizes well to unheard target sounds.

音视频导航条件融合空间感知多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。