提升社交媒体情感分析的视觉质量与时间建模,让图像不再拖后腿。
Fine-Grained Visual Preprocessing and Dual-Stream Temporal Modeling for Multimodal Sentiment Analysis on Social Media

- 设计七阶段降噪流程,精准提取人脸与嘴部动作特征。
- 结合静态图像与光流运动信息,视觉准确率提升至82.58%。
- 适合数据有限场景下的多模态情感分析研究者使用。
由于原始视频噪声和时间建模不足,多模态情感分析常依赖文本。本研究基于CH-SIMS v2.0S数据集,提出三项改进:NAPS流水线——整合人脸追踪、身份嵌入与归一化唇动分析的七阶段系统,降低视觉噪声;DS-TANet——融合EfficientNetB2静态流、RAFT光流运动流、运动引导注意力与Bi-GRU时间建模;DS-TAFNet——通过拼接融合视觉与MacBERT-Base文本表示。在NAPS支持下,静态视觉基线达80.98%宏平均F1,接近文本基线的80.55%;DS-TANet将视觉宏平均F1提升至82.58%;DS-TAFNet实现87.49%准确率与87.48%宏平均F1。结果表明,在数据受限条件下,优化视觉输入质量与时间表征比复杂融合更有效。
原文摘要 · Abstract (English)
Multimodal sentiment analysis often remains text-dominant due to raw-video noise and insufficient temporal modeling. Using CH-SIMS v2.0S, this study proposes three improvements: the NAPS pipeline---a seven-stage system integrating face tracking,identity embedding, and normalized lip-motion analysis to reduce visual noise;DS-TANet, combining an EfficientNetB2 static stream, RAFT optical-flow motion stream, motion-guided attention, and Bi-GRU temporal modeling; and DS-TAFNet, fusing visual and MacBERT-Base textual representations via concatenation fusion. With NAPS, the static visual baseline achieves 80.98\% Macro F1, comparable to the text baseline of 80.55\%; DS-TANet improves visual Macro F1 to 82.58\%;and DS-TAFNet achieves 87.49\% accuracy and 87.48\% Macro F1. These results demonstrate that improving visual input quality and temporal representation is more effective than increasing fusion complexity under limited-data conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。