arXiv:2508.17497cs.LGcs.AI2025-08被引 1

让图像文本表示随语义关系动态变化,提升多模态理解能力

Multimodal Representation Learning Conditioned on Semantic Relations

  • 以语义关系为条件生成不同嵌入,同一内容在不同关系中表现不同
  • 在多个数据集上零样本、微调和跨域任务中均优于基线模型
  • 适合需要细粒度关系理解的场景,如智能问答与复杂检索

多模态表示学习主要依赖对比模型(如CLIP),通过配对图文样本来对齐共享嵌入空间。这类方法通常为每条样本生成单一嵌入,不区分语义关系和上下文。然而,在真实应用中,样本间的相关性本质上依赖于语义关系,不同关系会强调多模态数据的不同方面。本文提出关系条件化多模态学习(RCML)框架,将语义关系显式作为表示学习的条件。不同于生成无关关系的嵌入,RCML根据自然语言描述的关系条件来调整表示,使同一样本在不同关系上下文中产生不同表达。该框架构建了关系感知的训练样本对,引入关系条件化模块以适应关系语义,并采用统一对比目标联合建模跨模态对齐与关系驱动的样本间结构。在多个数据集上的实验表明,RCML在零样本、微调及跨域设置下的检索与分类任务中持续优于强基线,验证了利用语义关系引导多模态表示学习的有效性。

原文摘要 · Abstract (English)

Multimodal representation learning has been largely driven by contrastive models such as CLIP, which learn a shared embedding space by aligning paired image-text samples. While effective for general-purpose representation learning, such models typically produce a single embedding per sample that is reused across different semantic relations and contexts. However, in many real-world applications, relevance between samples is inherently relation-dependent, with different semantic relations emphasizing different aspects of multimodal data. In this work, we propose Relation-Conditioned Multimodal Learning (RCML), a framework that treats semantic relations as explicit conditions of multimodal representation learning. Rather than producing relation-agnostic embeddings, RCML learns representations conditioned on natural-language relation descriptions, allowing the same sample to be represented differently under different relational contexts. The framework constructs relation-aware training pairs, introduces a relation-conditioned module to adapt embeddings to relation semantics, and employs a unified contrastive objective to jointly model cross-modal alignment and relation-induced inter-sample structure. Experiments on multiple datasets show that RCML consistently outperforms strong baselines on retrieval and classification tasks in zero-shot, fine-tuned, and out-of-domain settings, highlighting the effectiveness of leveraging semantic relations to guide multimodal representation learning.

多模态关系建模表示学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。