融合方言语音与画作,提升古诗情感分析准确率
Picturized and Recited with Dialects: A Multimodal Chinese Representation Framework for Sentiment Analysis of Classical Chinese Poetry
- 引入方言语音与绘画视觉特征,增强古诗多模态表征
- 在两个数据集上准确率提升至少2.51%,宏平均F1提升1.63%
- 适合对古诗、多模态中文表示感兴趣的学者与开发者
古典汉诗是中华文学的重要组成部分,蕴含深厚情感。现有研究多基于文本语义进行情感分析,忽视了诗歌特有的韵律与视觉特征,尤其因其常伴随诵读与中国画呈现。本文提出一种融合方言的多模态框架,从诗句中提取句级音频特征,并引入多种方言音频,以保留区域性的古汉语发音特征,丰富语音表征。同时生成句级视觉特征,通过多模态对比表示学习,将增强后的文本特征与音频、视觉特征融合。该框架在两个公开数据集上优于现有最先进方法,准确率提升至少2.51%,宏平均F1提升1.63%。代码已开源,为通用多模态中文表征研究提供参考。
原文摘要 · Abstract (English)
Classical Chinese poetry is a vital and enduring part of Chinese literature, conveying profound emotional resonance. Existing studies analyze sentiment based on textual meanings, overlooking the unique rhythmic and visual features inherent in poetry,especially since it is often recited and accompanied by Chinese paintings. In this work, we propose a dialect-enhanced multimodal framework for classical Chinese poetry sentiment analysis. We extract sentence-level audio features from the poetry and incorporate audio from multiple dialects,which may retain regional ancient Chinese phonetic features, enriching the phonetic representation. Additionally, we generate sentence-level visual features, and the multimodal features are fused with textual features enhanced by LLM translation through multimodal contrastive representation learning. Our framework outperforms state-of-the-art methods on two public datasets, achieving at least 2.51% improvement in accuracy and 1.63% in macro F1. We open-source the code to facilitate research in this area and provide insights for general multimodal Chinese representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。