arXiv:2510.11175cs.CV2025-10被引 3

通过迭代构建原型,有效抑制跨模态对齐中的风格干扰。

Reliable Cross-modal Alignment via Prototype Iterative Construction

  • 基于特征列语义概率加权,动态分离语义与风格信息。
  • 在多个基准上提升5.2%至14.1%,超越当前最优方法。
  • 适合需要高鲁棒性跨模态对齐的研究者或应用开发者。

跨模态对齐旨在弥合不同模态间的语义鸿沟。现有方法隐式假设嵌入仅包含语义信息,忽略非语义信息(如风格)的干扰,导致信息偏差或丢失。本文提出PICO框架,通过量化每个特征列的语义概率作为交互权重,抑制风格干扰。关键在于性能反馈驱动的原型迭代构建机制,理论上证明该机制能为提升性能更显著的原型分配更高权重。大量实验表明,PICO在多个基准和模型架构上均优于现有方法,性能提升达5.2%~14.1%。

原文摘要 · Abstract (English)

Cross-modal alignment is an important multi-modal task, aiming to bridge the semantic gap between different modalities. The most reliable fundamention for achieving this objective lies in the semantic consistency between matched pairs. Conventional methods implicitly assume embeddings contain solely semantic information, ignoring the impact of non-semantic information during alignment, which inevitably leads to information bias or even loss. These non-semantic information primarily manifest as stylistic variations in the data, which we formally define as style information. An intuitive approach is to separate style from semantics, aligning only the semantic information. However, most existing methods distinguish them based on feature columns, which cannot represent the complex coupling relationship between semantic and style information. In this paper, we propose PICO, a novel framework for suppressing style interference during embedding interaction. Specifically, we quantify the probability of each feature column representing semantic information, and regard it as the weight during the embedding interaction. To ensure the reliability of the semantic probability, we propose a prototype iterative construction method. The key operation of this method is a performance feedback-based weighting function, and we have theoretically proven that the function can assign higher weight to prototypes that bring higher performance improvements. Extensive experiments on various benchmarks and model backbones demonstrate the superiority of PICO, outperforming state-of-the-art methods by 5.2\%-14.1\%.

跨模态对齐原型学习风格无关嵌入优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。