arXiv:2602.03817cs.SDcs.AI2026-02

根据输入自适应融合音频与时空信息,提升生物声学分类准确率。

Adaptive Evidence Weighting for Audio-Spatiotemporal Fusion

  • 设计可学习的门控机制,动态评估上下文信息可靠性。
  • 在CBI和BirdSet数据集上超越固定权重融合与纯音频模型。
  • 轻量级可解释框架,适合对结果可信度有要求的应用场景。

许多机器学习系统需整合多种证据进行预测,但不同证据在不同输入下的可靠性与信息量差异显著。在生物声学分类中,物种识别既依赖音频信号,也依赖位置与季节等时空上下文;尽管贝叶斯推理支持乘法融合,但实际中通常仅有判别性预测器而非校准的生成模型。本文提出融合独立条件假设(FINCH)框架,将预训练音频分类器与结构化时空预测器结合,通过每样本的门控函数,基于不确定性与信息量统计估算上下文信息的可靠性。该融合方法包含纯音频分类器作为特例,显式限制上下文证据影响,形成风险可控、可解释且具备音频回退能力的假设类。在多个基准测试中,FINCH持续优于固定权重融合与纯音频基线,即使上下文信息孤立时较弱,仍提升鲁棒性与误差权衡。在CBI上达到当前最优,在BirdSet多个子集上取得竞争力或更优表现。代码已公开:https://anonymous.4open.science/r/birdnoise-85CD/README.md

原文摘要 · Abstract (English)

Many machine learning systems have access to multiple sources of evidence for the same prediction target, yet these sources often differ in reliability and informativeness across inputs. In bioacoustic classification, species identity may be inferred both from the acoustic signal and from spatiotemporal context such as location and season; while Bayesian inference motivates multiplicative evidence combination, in practice we typically only have access to discriminative predictors rather than calibrated generative models. We introduce \textbf{F}usion under \textbf{IN}dependent \textbf{C}onditional \textbf{H}ypotheses (\textbf{FINCH}), an adaptive log-linear evidence fusion framework that integrates a pre-trained audio classifier with a structured spatiotemporal predictor. FINCH learns a per-sample gating function that estimates the reliability of contextual information from uncertainty and informativeness statistics. The resulting fusion family \emph{contains} the audio-only classifier as a special case and explicitly bounds the influence of contextual evidence, yielding a risk-contained hypothesis class with an interpretable audio-only fallback. Across benchmarks, FINCH consistently outperforms fixed-weight fusion and audio-only baselines, improving robustness and error trade-offs even when contextual information is weak in isolation. We achieve state-of-the-art performance on CBI and competitive or improved performance on several subsets of BirdSet using a lightweight, interpretable, evidence-based approach. Code is available: \texttt{\href{https://anonymous.4open.science/r/birdnoise-85CD/README.md}{anonymous-repository}}

音频处理多模态融合生物声学可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。