提出连续情感评分方法,更好分析文学作品中的细微情感变化。
Continuous sentiment scores for literary and multilingual contexts
- 基于概念向量投影的连续情感评分,适配多语言文学文本
- 在英、丹文文本上超越现有工具,结果接近人工标注分布
- 适合需要精细情感分析的文学研究与跨语言比较
情感分析广泛用于量化文本情绪,但应用于文学文本时面临隐喻语言、风格模糊及情感唤起策略等独特挑战。传统词典工具表现不佳,尤其在低资源语言中;而基于Transformer的模型虽有潜力,通常输出粗粒度分类标签,限制了细粒度分析。本文提出一种基于概念向量投影的连续情感评分方法,训练于多语言文学数据,能更有效捕捉跨体裁、语言和历史时期的细微情感表达。该方法在英文和丹麦文文本上优于现有工具,其情感得分分布与人工评分高度一致,支持更精准的文学情感演变建模。
原文摘要 · Abstract (English)
Sentiment Analysis is widely used to quantify sentiment in text, but its application to literary texts poses unique challenges due to figurative language, stylistic ambiguity, as well as sentiment evocation strategies. Traditional dictionary-based tools often underperform, especially for low-resource languages, and transformer models, while promising, typically output coarse categorical labels that limit fine-grained analysis. We introduce a novel continuous sentiment scoring method based on concept vector projection, trained on multilingual literary data, which more effectively captures nuanced sentiment expressions across genres, languages, and historical periods. Our approach outperforms existing tools on English and Danish texts, producing sentiment scores whose distribution closely matches human ratings, enabling more accurate analysis and sentiment arc modeling in literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。