arXiv:2602.03570cs.LG2026-02

提出不对称层级锚定机制,提升音视频跨模态迁移的鲁棒性。

Asymmetric Hierarchical Anchoring for Robust Audio-Visual Cross-Modal Generalization

  • 设计不对称层级锚点,引导信息单向分配,避免模态间语义混淆。
  • 在AVE和AVVP数据集上,跨模态迁移性能超越对称基线模型。
  • 适合关注音视频联合表征与跨模态泛化任务的研究者。

在跨模态泛化(CMG)框架下,音视频联合表征学习旨在通过统一离散表示空间,将带标签源模态的知识迁移到无标签目标模态。现有对称框架常因缺乏结构归纳偏置,导致模态间语义泄漏。本文提出不对称层级锚定(AHA),通过在共享层次中设定结构化语义锚点,强制定向信息分配。具体实现中,利用音频残差向量量化(RVQ)生成的层次化离散表示,指导视频特征蒸馏至共享语义空间。为保障表征纯净性,用基于梯度反向学习(GRL)的对抗解耦器替代脆弱的互信息估计器,显式抑制模态特异性分支中的语义泄漏,并引入局部滑动对齐(LSA)促进模态间细粒度时间对齐。在AVE和AVVP基准上的大量实验表明,AHA在跨模态迁移中持续优于对称基线。额外的说话人脸解耦实验进一步验证,所学表征具备更优的语义一致性和解耦性,证明该框架具有更广适用性。

原文摘要 · Abstract (English)

Audio-visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space. Existing symmetric frameworks often suffer from information allocation ambiguity, where the absence of structural inductive bias leads to semantic-specific leakage across modalities. We propose Asymmetric Hierarchical Anchoring (AHA), which enforces directional information allocation by designating a structured semantic anchor within a shared hierarchy. In our instantiation, we exploit the hierarchical discrete representations induced by audio Residual Vector Quantization (RVQ) to guide video feature distillation into a shared semantic space. To ensure representational purity, we replace fragile mutual information estimators with a GRL-based adversarial decoupler that explicitly suppresses semantic leakage in modality-specific branches, and introduce Local Sliding Alignment (LSA) to encourage fine-grained temporal alignment across modalities. Extensive experiments on AVE and AVVP benchmarks demonstrate that AHA consistently outperforms symmetric baselines in cross-modal transfer. Additional analyses on talking-face disentanglement experiment further validate that the learned representations exhibit improved semantic consistency and disentanglement, indicating the broader applicability of the proposed framework.

跨模态学习音视频表征学习解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。