arXiv:2608.25520cs.CV2026-08中稿 · the 9th Chinese Co…

解决音视频不同步下的鸟类细粒度识别难题,提升弱匹配场景下的识别准确率。

Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark

论文配图:Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark
图 1 · 摘自论文原文
  • 用光流引导捕捉动态视觉特征,抑制背景干扰,增强视频表征。
  • 设计不确定性感知的自适应融合机制,在非严格配对下仍保持高精度。
  • 构建新基准BirdPro,含194种鸟的1.2万段音视频数据,支持弱对应研究。

音频-视觉跨模态细粒度图像分类旨在联合利用视觉与听觉信息识别细粒度类别。然而,非对称跨模态场景下的细粒度分类研究尚不充分,其中音视频对未严格同步,甚至可能对应不同个体或时间点,这种弱且模糊的跨模态对应关系给有效表征学习和模态对齐带来巨大挑战。为此,我们提出ACF-Net,一种基于光流引导的非对称音视频细粒度学习框架。ACF-Net包含两个核心模块:光流引导运动建模(OFGM)与非对称跨模态自适应融合(ACAF)。OFGM捕捉运动敏感的视觉线索并抑制无关背景干扰,从而增强视频中的判别性动态表征。ACAF在弱匹配音视频对下估计模态可靠性,并执行不确定性感知的自适应融合,提升类别级识别鲁棒性。为支持非对称跨模态细粒度分类研究,我们进一步构建了面向鸟类的BirdPro新基准数据集,因其现有数据集普遍缺乏在非严格时间与实例对应条件下的大规模类别级音视频关联。BirdPro包含1,919条音频记录和11,965段视频,覆盖194种鸟类。大量实验表明,相较于代表性基线方法,ACF-Net在融合与错位设置下分别领先最强基线2.97%和1.92%。

原文摘要 · Abstract (English)

Audio-visual cross-modal Fine-Grained Visual Categorization (FGVC) aims to identify fine-grained categories by jointly leveraging visual and auditory information. However, FGVC under asymmetric cross-modal scenarios has received limited attention, where paired video and audio are not strictly synchronized and may not even correspond to the same individual or moment. Such weak and ambiguous cross-modal correspondence poses substantial challenges to effective representation learning and modality alignment. To address these issues, we propose ACF-Net, a novel optical flow-guided framework for asymmetric audio-visual fine-grained learning. ACF-Net consists of two key modules: Optical Flow-Guided Motion (OFGM) and Asymmetric CrossModal Adaptive Fusion (ACAF). OFGM captures motion-sensitive visual cues and suppresses irrelevant background interference, thereby enhancing discriminative dynamic representations in videos. ACAF estimates modality reliability under weakly matched audio-video pairs and performs uncertainty-aware adaptive fusion to improve category-level recognition robustness. To support research on asymmetric cross-modal FGVC, we further construct BirdPro, a new bird-oriented audio-visual benchmark, since existing datasets often lack large-scale category-level audio-video associations under non-strict temporal and instance correspondence. BirdPro contains 1,919 audio recordings and 11,965 videos covering 194 bird species. Extensive experiments show that ACF-Net achieves the best results compared with representative baseline methods, outperforming the strongest baselines by 2.97% and 1.92% in the fused and mismatched settings, respectively.

细粒度识别音视频融合跨模态学习鸟类识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。