分析斯里兰卡旅游点评,发现近两成评分与文字情感不符
Fault of Our Stars: Behavioral Drivers of Rating-Sentiment Incongruence
- 用Transformer模型独立分析文本情感,对比用户评分
- 18.6%点评存在评分与文字情感不一致,博物馆最明显
- 评分偏差受景点类型、评论长度、用户经验等影响
当人们在线分享体验时,常同时给出星级评分和文字评论。尽管评分常被用作文本情感的弱标签,但二者是否一致却少被质疑。本研究以2010至2023年斯里兰卡旅游景点评论数据(共16,156条)为样本,采用基于Transformer的文本情感分析管道,独立提取文本情感,发现18.6%的评论存在评分-情感不一致现象,可分为六类模式,其中保守评分者和强制给五星行为占主导。不同场所类型差异显著,博物馆不一致率最高。统计检验、逻辑回归、随机森林及SHAP分析表明,场所类型、用户专业度、评论长度与时间因素是导致评分与文本情感偏离的关键因素。研究证明,星级评分不可替代文本情感,需在自然语言处理任务中验证其作为真实标签的可靠性。
原文摘要 · Abstract (English)
When people share experiences online, they often express thoughts in two ways: a star rating and a written review. In sentiment analysis, ratings are widely used as convenient weak labels for textual sentiment, yet whether the two actually agree is rarely questioned. This study investigates sentiment-rating incongruence, where the sentiment expressed in review text differs from the sentiment implied by the assigned star rating, in Sri Lankan tourism attraction reviews. A dataset of 16,156 reviews from 2010 to 2023 is analyzed using a transformer-based sentiment pipeline that derives textual sentiment independently of assigned ratings. Incongruence occurs in 18.6% of reviews and falls into six directional patterns, with Conservative Rater and Obligatory 5-Star behaviors accounting for the majority of mismatches. Prevalence also varies across venue types, with museums showing the highest rates. Statistical tests, logistic regression, Random Forest, and SHAP analysis identify venue type, reviewer expertise, review length, and temporal factors as contributors to rating-text divergence. Overall, this study demonstrates that star ratings are not interchangeable with textual sentiment and should be validated before being treated as ground-truth labels in NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。