用对比学习生成虚拟弹幕特征,提升跨语言情感分析效果
Enhancing Multimodal Affective Analysis with Learned Live Comment Features

- 通过对比学习训练视频编码器生成合成弹幕特征
- 在中英文多任务上均显著超越现有方法
- 适合做多模态情感分析与跨语言内容理解的研究者
实时弹幕(Danmaku)是用户在观看视频时同步发送的评论,直接叠加于视频之上,能捕捉观众即时情绪。尽管已有研究尝试利用弹幕进行情感分析,但因不同平台弹幕数据稀少而受限。为此,我们构建了涵盖中英文多样视频类型的LCAffect数据集,包含广泛情绪表现。基于该数据集,采用对比学习训练视频编码器,生成用于增强多模态情感分析的合成弹幕特征。在英语和中文环境下,针对情感、情绪识别及反讽检测等任务的全面实验表明,该方法显著优于当前最优模型。
原文摘要 · Abstract (English)
Live comments, also known as Danmaku, are user-generated messages that are synchronized with video content. These comments overlay directly onto streaming videos, capturing viewer emotions and reactions in real-time. While prior work has leveraged live comments in affective analysis, its use has been limited due to the relative rarity of live comments across different video platforms. To address this, we first construct the Live Comment for Affective Analysis (LCAffect) dataset which contains live comments for English and Chinese videos spanning diverse genres that elicit a wide spectrum of emotions. Then, using this dataset, we use contrastive learning to train a video encoder to produce synthetic live comment features for enhanced multimodal affective content analysis. Through comprehensive experimentation on a wide range of affective analysis tasks (sentiment, emotion recognition, and sarcasm detection) in both English and Chinese, we demonstrate that these synthetic live comment features significantly improve performance over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。