arXiv:2605.12145cs.CV2026-05

用离散表示统一多模态信息,兼顾特异性与泛化性。

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations

论文配图:Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations
图 1 · 摘自论文原文
  • 通过索引级对齐实现跨模态代码本语义一致
  • 在多个任务上超越现有方法,最高提升6.3%
  • 适合需要强泛化能力的多模态系统设计

多模态学习旨在融合不同感官来源的信息,但现有方法难以平衡跨模态泛化性与模态特异性。连续(隐式)方法保留精细先验,但泛化困难;离散(显式)方法虽强化共享原型,却牺牲模态特异性。本文提出 CoDAAR(跨模态离散对齐与重建),通过模态特定代码本间的索引级对齐,建立跨模态语义共识,既保留模态独特结构,又在统一离散空间中实现可泛化的多模态表示。CoDAAR 结合两种互补机制:离散时间对齐(DTA)实现细粒度时间量化,级联语义对齐(CSA)促进渐进式跨模态语义一致性。二者共同构建无竞争的统一表示空间。在成对多模态序列上使用自监督重建目标训练,CoDAAR 在事件分类、定位、视频分割及跨数据集迁移等跨模态泛化基准上表现优异,达到当前最优性能,为离散且可泛化的多模态表征学习树立新范式。

原文摘要 · Abstract (English)

Multimodal learning seeks to integrate information across diverse sensory sources, yet current approaches struggle to balance cross-modal generalizability with modality-specific structure. Continuous (implicit) methods preserve fine-grained priors but render generalization challenging, while discrete (explicit) approaches enforce shared prototypes at the expense of modality specificity. We introduce CoDAAR (Cross-modal Discrete Alignment And Reconstruction), a novel framework that resolves this long-standing trade-off by establishing semantic consensus across modality-specific codebooks through index-level alignment. This design uniquely allows CoDAAR to preserve modality-unique structures while achieving generalizable cross-modal representations within a unified discrete space. CoDAAR combines two complementary mechanisms: Discrete Temporal Alignment (DTA), which enables fine-grained temporal quantization, and Cascading Semantic Alignment (CSA), which promotes progressive cross-modal semantic agreement. Together, they establish a competition-free unified representation space. Trained with self-supervised reconstruction objectives on paired multimodal sequences, CoDAAR demonstrates robust cross-modal and cross-domain generalization. Across Cross-Modal Generalization benchmarks, including event classification, localization, video segmentation, and cross-dataset transfer, CoDAAR achieves state-of-the-art performance, establishing a new paradigm for discrete and generalizable multimodal representation learning.

多模态学习离散表示泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。