不依赖目标语言数据,实现多语言新闻情感分析。
Evaluating and explaining training strategies for zero-shot cross-lingual news sentiment analysis

- 用上下文学习与新任务目标POA提升跨语言情感分类。
- 上下文学习表现最佳,POA计算开销低且效果接近。
- 语义相似性比语言相似性更影响跨语言迁移效果。
我们研究零样本跨语言新闻情感检测,旨在开发无需目标语言训练数据即可部署的鲁棒情感分类器。针对若干资源较少的语言构建了新的评估数据集,并实验了多种方法,包括机器翻译、大语言模型的上下文学习,以及多种中间训练策略,其中提出一种新任务目标POA,利用段落级信息。结果表明性能显著优于现有方法,上下文学习通常表现最佳,但新提出的POA方法在计算开销大幅降低的情况下仍具竞争力。此外,我们发现语言相似性本身不足以预测跨语言迁移的成功,而语义内容和结构的相似性同样关键。
原文摘要 · Abstract (English)
We investigate zero-shot cross-lingual news sentiment detection, aiming to develop robust sentiment classifiers that can be deployed across multiple languages without target-language training data. We introduce novel evaluation datasets in several less-resourced languages, and experiment with a range of approaches including the use of machine translation; in-context learning with large language models; and various intermediate training regimes including a novel task objective, POA, that leverages paragraph-level information. Our results demonstrate significant improvements over the state of the art, with in-context learning generally giving the best performance, but with the novel POA approach giving a competitive alternative with much lower computational overhead. We also show that language similarity is not in itself sufficient for predicting the success of cross-lingual transfer, but that similarity in semantic content and structure can be equally important.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。