arXiv:2507.00926cs.MMcs.LG2025-07被引 3

用多模态融合提升社交媒体内容热度预测准确率。

HyperFusion: Hierarchical Multimodal Ensemble Learning for Social Media Popularity Prediction

  • 分层融合视觉、文本与时空行为特征,逐步整合信息
  • 在SMP挑战赛图像赛道获第三名,性能优于基线模型
  • 适合需要精准内容推荐的平台运营与营销团队

社交媒体热度预测对内容优化、营销策略和用户参与度提升至关重要。然而,由于视觉、文本、时间与用户行为因素的复杂交互,预测仍具挑战。本文提出HyperFusion,一种分层多模态集成学习框架。该框架采用三层融合结构,逐步整合来自CLIP编码器的视觉表征、Transformer模型的文本嵌入,以及时间-空间元数据与用户特征。通过结合CatBoost、TabNet与自定义多层感知机实现分层集成。针对标注数据有限问题,设计两阶段训练法,包含伪标签与迭代优化。引入新颖的跨模态相似性度量与分层聚类特征以捕捉模态间依赖关系。实验表明,HyperFusion在SMP挑战赛数据集上表现优异,团队在SMP Challenge 2025(Image Track)中获第三名。源代码已公开于https://anonymous.4open.science/r/SMPDImage。

原文摘要 · Abstract (English)

Social media popularity prediction plays a crucial role in content optimization, marketing strategies, and user engagement enhancement across digital platforms. However, predicting post popularity remains challenging due to the complex interplay between visual, textual, temporal, and user behavioral factors. This paper presents HyperFusion, a hierarchical multimodal ensemble learning framework for social media popularity prediction. Our approach employs a three-tier fusion architecture that progressively integrates features across abstraction levels: visual representations from CLIP encoders, textual embeddings from transformer models, and temporal-spatial metadata with user characteristics. The framework implements a hierarchical ensemble strategy combining CatBoost, TabNet, and custom multi-layer perceptrons. To address limited labeled data, we propose a two-stage training methodology with pseudo-labeling and iterative refinement. We introduce novel cross-modal similarity measures and hierarchical clustering features that capture inter-modal dependencies. Experimental results demonstrate that HyperFusion achieves competitive performance on the SMP challenge dataset. Our team achieved third place in the SMP Challenge 2025 (Image Track). The source code is available at https://anonymous.4open.science/r/SMPDImage.

多模态学习热度预测集成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。