不同情感分析工具结果差异大,其输出可被反推,说明存在算法偏见。
You Shall Know a Tool by the Traces it Leaves: The Predictability of Sentiment Analysis Tools
- 通过分析工具输出预测使用了哪个工具,揭示算法偏见
- 在英文学科数据上平均F1达0.89,证明预测有效
- 提醒研究者勿轻信情感标注,需更系统评估
若情感分析工具是有效分类器,应在不同语料和语言上给出一致结果。然而,与以往研究一致,我们发现不同工具对同一数据集的判断存在分歧。进一步地,我们证明仅凭情感分析结果即可预测出所用工具,揭示其内在算法偏见。基于英文、德文、法文的Twitter、Wikipedia及新闻语料,我们的分类器在英文学科数据上实现平均F1分数0.89。因此,我们警告不应将情感标注视为可靠事实,并呼吁开展更系统、更广泛的自然语言处理评估研究。
原文摘要 · Abstract (English)
If sentiment analysis tools were valid classifiers, one would expect them to provide comparable results for sentiment classification on different kinds of corpora and for different languages. In line with results of previous studies we show that sentiment analysis tools disagree on the same dataset. Going beyond previous studies we show that the sentiment tool used for sentiment annotation can even be predicted from its outcome, revealing an algorithmic bias of sentiment analysis. Based on Twitter, Wikipedia and different news corpora from the English, German and French languages, our classifiers separate sentiment tools with an averaged F1-score of 0.89 (for the English corpora). We therefore warn against taking sentiment annotations as face value and argue for the need of more and systematic NLP evaluation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。