arXiv:2603.29820cs.SD2026-03

用视觉信息将单声道视频转为沉浸式立体声,提升听觉空间感。

SIREN: Spatially-Informed Reconstruction of Binaural Audio with Vision

  • 基于视觉的双通道注意力机制,自动学习左右耳音频映射。
  • 在两个数据集上显著提升时间-频率与相位指标,信噪比表现优秀。
  • 无需人工标注,可无缝接入现有音视频处理流程,适合多场景应用。

立体声音频能提供关键的空间听觉线索,但受拍摄条件限制,多数消费级视频为单声道。本文提出 SIREN,一种基于视觉引导的单声道转立体声框架,显式预测左右声道。采用 ViT 编码器,通过双头自注意力生成共享场景图与端到端左右耳注意力,替代手工设计的掩码。引入柔和且渐进的时空先验,以早期引导左右声道定位;再通过两阶段、置信度加权的波形域融合(基于单声道重建和双耳相位一致性),有效抑制多裁剪重叠窗口聚合时的串扰。在 FAIR-Play 与 MUSIC-Stereo 数据集上的评估显示,SIREN 在时间-频率及相位敏感指标上均取得一致提升,且信噪比表现具有竞争力。该设计模块化且通用,无需任务特定标注,可集成至标准音视频处理流程。

原文摘要 · Abstract (English)

Binaural audio delivers spatial cues essential for immersion, yet most consumer videos are monaural due to capture constraints. We introduce SIREN, a visually guided mono to binaural framework that explicitly predicts left and right channels. A ViT-based encoder learns dual-head self-attention to produce a shared scene map and end-to-end L/R attention, replacing hand-crafted masks. A soft, annealed spatial prior gently biases early L/R grounding, and a two-stage, confidence-weighted waveform-domain fusion (guided by mono reconstruction and interaural phase consistency) suppresses crosstalk when aggregating multi-crop and overlapping windows. Evaluated on FAIR-Play and MUSIC-Stereo, SIREN yields consistent gains on time-frequency and phase-sensitive metrics with competitive SNR. The design is modular and generic, requires no task-specific annotations, and integrates with standard audio-visual pipelines.

立体声生成视觉引导音频增强ViT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。