arXiv:2507.04508cs.CL2025-07被引 1

用轻量适配器共享机制提升图文讽刺检测效率

Adapter-state Sharing CLIP for Parameter-efficient Multimodal Sarcasm Detection

  • 仅在高层插入适配器,保留底层单模态特征
  • 文本适配器引导视觉适配器,跨模态学习更高效
  • 参数少却超越主流方法,适合资源受限场景

社交媒体上多模态图文讽刺日益普遍,给意见挖掘系统带来挑战。现有方法依赖大模型全量微调,难以在资源受限环境下应用。尽管参数高效微调(PEFT)有潜力,但在讽刺检测等复杂任务上表现不佳。本文提出AdS-CLIP,在CLIP基础上仅在高层插入适配器,保留底层单模态表示,并引入适配器状态共享机制:文本适配器指导视觉适配器,促进高层跨模态学习。在两个公开数据集上的实验表明,AdS-CLIP不仅优于标准PEFT方法,也超过现有多模态基线,且可训练参数显著减少。

原文摘要 · Abstract (English)

The growing prevalence of multimodal image-text sarcasm on social media poses challenges for opinion mining systems. Existing approaches rely on full fine-tuning of large models, making them unsuitable to adapt under resource-constrained settings. While recent parameter-efficient fine-tuning (PEFT) methods offer promise, their off-the-shelf use underperforms on complex tasks like sarcasm detection. We propose AdS-CLIP (Adapter-state Sharing in CLIP), a lightweight framework built on CLIP that inserts adapters only in the upper layers to preserve low-level unimodal representations in the lower layers and introduces a novel adapter-state sharing mechanism, where textual adapters guide visual ones to promote efficient cross-modal learning in the upper layers. Experiments on two public benchmarks demonstrate that AdS-CLIP not only outperforms standard PEFT methods but also existing multimodal baselines with significantly fewer trainable parameters.

多模态讽刺检测参数高效CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。