arXiv:2604.04496cs.CV2026-04NeurIPS被引 8

提出用范畴论构建跨模态统一关系表征,无需训练即可对齐不同模型。

The Indra Representation Hypothesis for Multimodal Alignment

  • 基于范畴论的Yoneda嵌入,将样本间关系建模为唯一且保结构的表征。
  • 在视觉、语言、音频跨模态任务中,提升模型对齐效果与鲁棒性。
  • 适用于不需微调的多模态对齐,适合希望统一不同模型表征的研究者。

近期研究发现,无论架构、训练目标或数据模态如何,单模态基础模型倾向于学习收敛的表征,但这些表征仅是样本的独立抽象,表达能力有限。本文提出‘因陀罗表征假说’,受哲学隐喻‘因陀罗之网’启发,认为单模态模型的表征正收敛于现实背后共享的关系结构。我们通过范畴论中的V-丰富Yoneda嵌入形式化该假说,将因陀罗表征定义为每个样本相对于其他样本的关系谱。该表征在给定代价函数下具有唯一性、完备性与结构保持性。我们使用角度距离实例化该表征,并在涉及视觉、语言和音频的跨模型与跨模态场景中进行评估。大量实验表明,因陀罗表征能持续增强跨架构与跨模态的鲁棒性与对齐性能,提供了一种理论严谨且实用的免训练对齐框架。代码已开源:https://github.com/Jianglin954/Indra。

原文摘要 · Abstract (English)

Recent studies have uncovered an interesting phenomenon: unimodal foundation models tend to learn convergent representations, regardless of differences in architecture, training objectives, or data modalities. However, these representations are essentially internal abstractions of samples that characterize samples independently, leading to limited expressiveness. In this paper, we propose The Indra Representation Hypothesis, inspired by the philosophical metaphor of Indra's Net. We argue that representations from unimodal foundation models are converging to implicitly reflect a shared relational structure underlying reality, akin to the relational ontology of Indra's Net. We formalize this hypothesis using the V-enriched Yoneda embedding from category theory, defining the Indra representation as a relational profile of each sample with respect to others. This formulation is shown to be unique, complete, and structure-preserving under a given cost function. We instantiate the Indra representation using angular distance and evaluate it in cross-model and cross-modal scenarios involving vision, language, and audio. Extensive experiments demonstrate that Indra representations consistently enhance robustness and alignment across architectures and modalities, providing a theoretically grounded and practical framework for training-free alignment of unimodal foundation models. Our code is available at https://github.com/Jianglin954/Indra.

多模态对齐范畴论表征学习免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。