arXiv:2507.04468cs.CL2025-07被引 1

针对少样本场景,用分模态深度提示提升跨模态讽刺检测效果

Dual Modality-Aware Gated Prompt Tuning for Few-Shot Multimodal Sarcasm Detection

  • 为文本和图像分别设计分层深度提示,增强特征学习
  • 在少样本和极低资源下超越现有方法,跨数据集泛化能力强
  • 适合社交媒体讽刺内容分析,尤其标注数据稀缺的场景

社交媒体上多模态内容的广泛使用,促使人们更迫切需要有效的讽刺检测以提升意见挖掘效果。然而,现有模型严重依赖大规模标注数据,在真实场景中标签稀缺时表现不佳。为此,我们提出 DMDP(Deep Modality-Disentangled Prompt Tuning),一种面向少样本多模态讽刺检测的新框架。不同于以往在各模态使用统一浅层提示的方法,DMDP 采用门控的、模态特定的深层提示,分别注入文本和视觉编码器的多个层次,实现层级特征学习,更好捕捉多样化的讽刺类型。为增强模态内学习,引入跨层提示共享机制,聚合低层与高层语义线索;同时,设计跨模态提示对齐模块,促进图像与文本表征间的细粒度交互,提升对微妙讽刺意图的识别能力。在两个公开数据集上的实验表明,DMDP 在少样本及极低资源设置下均表现卓越。跨数据集评估进一步验证其良好的域泛化能力,持续优于基线方法。

原文摘要 · Abstract (English)

The widespread use of multimodal content on social media has heightened the need for effective sarcasm detection to improve opinion mining. However, existing models rely heavily on large annotated datasets, making them less suitable for real-world scenarios where labeled data is scarce. This motivates the need to explore the problem in a few-shot setting. To this end, we introduce DMDP (Deep Modality-Disentangled Prompt Tuning), a novel framework for few-shot multimodal sarcasm detection. Unlike prior methods that use shallow, unified prompts across modalities, DMDP employs gated, modality-specific deep prompts for text and visual encoders. These prompts are injected across multiple layers to enable hierarchical feature learning and better capture diverse sarcasm types. To enhance intra-modal learning, we incorporate a prompt-sharing mechanism across layers, allowing the model to aggregate both low-level and high-level semantic cues. Additionally, a cross-modal prompt alignment module enables nuanced interactions between image and text representations, improving the model's ability to detect subtle sarcastic intent. Experiments on two public datasets demonstrate DMDP's superior performance in both few-shot and extremely low-resource settings. Further cross-dataset evaluations show that DMDP generalizes well across domains, consistently outperforming baseline methods.

多模态讽刺检测少样本学习提示调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。