通过原型增强与双粒度提示学习,提升社交媒体内容流行度预测效果
Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
- 构建分层原型并用对比学习加强图文对齐
- 双粒度提示学习实现细粒度类别建模,提升表示精度
- 适合做多模态社交内容分析的开发者参考
社交媒体流行度预测是一项复杂的多模态任务,需有效融合图像、文本和结构化信息。现有方法存在视觉-文本对齐不足,难以捕捉社交媒体数据中的跨内容关联与层次模式。为此,本文提出一种多类别框架,引入分层原型进行结构增强,并采用对比学习改善视觉-文本对齐。进一步设计特征增强型框架,集成双粒度提示学习与跨模态注意力机制,通过细粒度类别建模实现精确的多模态表征。实验结果在基准指标上达到领先水平,为多模态社交媒体分析建立新标准。
原文摘要 · Abstract (English)
Social Media Popularity Prediction is a complex multimodal task that requires effective integration of images, text, and structured information. However, current approaches suffer from inadequate visual-textual alignment and fail to capture the inherent cross-content correlations and hierarchical patterns in social media data. To overcome these limitations, we establish a multi-class framework , introducing hierarchical prototypes for structural enhancement and contrastive learning for improved vision-text alignment. Furthermore, we propose a feature-enhanced framework integrating dual-grained prompt learning and cross-modal attention mechanisms, achieving precise multimodal representation through fine-grained category modeling. Experimental results demonstrate state-of-the-art performance on benchmark metrics, establishing new reference standards for multimodal social media analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。