用少量目标语言数据提升跨语言情感分析效果
Few-shot Cross-lingual Aspect-Based Sentiment Analysis with Sequence-to-Sequence Models
- 用序列到序列模型,加入少量目标语言样本训练
- 仅10个标注样本就显著优于零样本,接近约束解码效果
- 1000个样本加英语数据可超越单语基线,适合低资源场景
方面级情感分析(ABSA)在英语中已获广泛关注,但在低资源语言中仍面临标注数据稀缺的挑战。现有跨语言ABSA方法常依赖外部翻译工具,且未充分利用少量目标语言样本的潜力。本文评估了在四个ABSA任务、六种目标语言及两种序列到序列模型上,向训练集添加少量目标语言样本的效果。结果表明,仅增加10个目标语言样本即可显著提升性能,优于零样本设置,并达到与约束解码相当的误差降低效果;此外,结合1000个目标语言样本与英语数据,甚至可超越单语基线。这些发现为低资源及特定领域中的跨语言ABSA提供了实用指导,因获取10个高质量标注样本既可行又高效。
原文摘要 · Abstract (English)
Aspect-based sentiment analysis (ABSA) has received substantial attention in English, yet challenges remain for low-resource languages due to the scarcity of labelled data. Current cross-lingual ABSA approaches often rely on external translation tools and overlook the potential benefits of incorporating a small number of target language examples into training. In this paper, we evaluate the effect of adding few-shot target language examples to the training set across four ABSA tasks, six target languages, and two sequence-to-sequence models. We show that adding as few as ten target language examples significantly improves performance over zero-shot settings and achieves a similar effect to constrained decoding in reducing prediction errors. Furthermore, we demonstrate that combining 1,000 target language examples with English data can even surpass monolingual baselines. These findings offer practical insights for improving cross-lingual ABSA in low-resource and domain-specific settings, as obtaining ten high-quality annotated examples is both feasible and highly effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。