arXiv:2509.16926cs.SDcs.AI2025-09中稿 · on Workshop on Det…被引 1

用置信度加权交叉注意力提升多通道音频同步精度。

Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment

  • 引入交叉注意力建模通道间时间依赖,融合置信度评分
  • 在生物声学数据上平均误差达0.30 MSE,较基线降低48%
  • 支持概率化对齐,适合需可靠性评估的音频同步场景

多通道音频同步是生物声学监测、空间音频系统和声源定位的关键。现有方法常无法处理非线性时钟漂移,且缺乏不确定性量化机制。传统方法如互相关和动态时间规整假设简单漂移模式,且不提供可靠性度量;近期深度学习模型通常将对齐视为二分类任务,忽视通道间依赖与不确定性估计。本文提出结合交叉注意力与置信度加权评分的方法,扩展BEATs编码器加入交叉注意力层以建模通道间时间关系,并设计基于完整预测分布的置信度加权评分函数,避免二值阈值。该方法在BioDCASE 2025 Task 1挑战中取得第一名,测试集平均MSE为0.30,优于基线深度学习模型的0.58。在单个数据集上,ARU数据实现0.14 MSE(降低77%),斑胸草雀数据实现0.45 MSE(降低18%)。该框架支持概率化时间对齐,超越点估计。虽在生物声学场景验证,但适用于需对齐置信度的各类多通道音频任务。代码已开源。

原文摘要 · Abstract (English)

Multi-channel audio alignment is a key requirement in bioacoustic monitoring, spatial audio systems, and acoustic localization. However, existing methods often struggle to address nonlinear clock drift and lack mechanisms for quantifying uncertainty. Traditional methods like Cross-correlation and Dynamic Time Warping assume simple drift patterns and provide no reliability measures. Meanwhile, recent deep learning models typically treat alignment as a binary classification task, overlooking inter-channel dependencies and uncertainty estimation. We introduce a method that combines cross-attention mechanisms with confidence-weighted scoring to improve multi-channel audio synchronization. We extend BEATs encoders with cross-attention layers to model temporal relationships between channels. We also develop a confidence-weighted scoring function that uses the full prediction distribution instead of binary thresholding. Our method achieved first place in the BioDCASE 2025 Task 1 challenge with 0.30 MSE average across test datasets, compared to 0.58 for the deep learning baseline. On individual datasets, we achieved 0.14 MSE on ARU data (77% reduction) and 0.45 MSE on zebra finch data (18% reduction). The framework supports probabilistic temporal alignment, moving beyond point estimates. While validated in a bioacoustic context, the approach is applicable to a broader range of multi-channel audio tasks where alignment confidence is critical. Code available on: https://github.com/Ragib-Amin-Nihal/BEATsCA

音频对齐交叉注意力置信度加权生物声学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。