arXiv:2508.13387cs.AI2025-08

让不同模态数据在统一空间对齐,提升跨模态泛化能力

SPANER: Shared Prompt Aligner for Multimodal Semantic Representation

  • 用共享提示词作为语义锚点,统一多模态输入表示
  • 在视觉-语言与音视频任务上实现竞争力的少样本检索性能
  • 架构可扩展,支持无缝加入音频等新模态

多模态参数高效微调(PEFT)近年显著提升了少样本检索等下游任务表现。然而,现有方法多关注任务特定提升,忽视多模态嵌入空间的结构。导致模态特异性表示相互孤立,限制跨模态泛化。本文提出共享提示对齐器(SPANER),一种无模态偏倚的PEFT框架,旨在将不同模态输入映射到统一语义空间。核心是共享提示机制,作为概念锚点,使语义相关的实例无论模态如何都能在空间中聚集。该设计天然可扩展,支持新增音频等模态而无需修改主架构。在视觉-语言与音视频基准上的实验表明,SPANER在保持高语义一致性的同时,实现了具有竞争力的少样本检索性能。结果强调,对齐嵌入结构比仅调优适配器权重对可扩展多模态学习更为重要。

原文摘要 · Abstract (English)

Recent advances in multimodal Parameter-Efficient Fine-Tuning (PEFT) have significantly improved performance on downstream tasks such as few-shot retrieval. However, most existing approaches focus on task-specific gains while neglecting the structure of the multimodal embedding space. As a result, modality-specific representations often remain isolated, limiting cross-modal generalisation. In this work, we introduce Shared Prompt AligNER (SPANER), a modality-agnostic PEFT framework designed to embed inputs from diverse modalities into a unified semantic space. At its core, SPANER employs a shared prompt mechanism that acts as a conceptual anchor, enabling semantically related instances to converge spatially regardless of modality. This shared prompt design is inherently extensible, supporting the seamless integration of additional modalities, such as audio, without altering the core architecture. Through comprehensive experiments across vision-language and audio-visual benchmarks, SPANER demonstrates competitive few-shot retrieval performance while preserving high semantic coherence in the learned embedding space. Our results highlight the importance of aligning embedding structures, rather than merely tuning adapter weights, for scalable multimodal learning.

多模态提示对齐参数高效微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。