跨语言情感分析新方法,让非英语文本也能精准识别观点。
Zero-Shot to Full-Resource: Cross-lingual Transfer Strategies for Aspect-Based Sentiment Analysis

- 用多语言迁移、混写和机器翻译提升非英语情感分析性能。
- 大模型微调在复杂任务中表现最佳,小模型适合混写策略。
- 新增德语数据集,推动非英语领域研究发展。
基于方面的情感分析(ABSA)可提取文本中针对特定方面的细粒度观点,但目前仍以英语为主。本文系统评估了七种语言(英语、德语、法语、荷兰语、俄语、西班牙语、捷克语)下最先进的ABSA方法,涵盖四个子任务(ACD、ACSA、TASD、ASQP)。在零资源、仅数据、全资源三种设置下,对比不同Transformer架构的表现,采用跨语言迁移、代码混写和机器翻译等策略。微调的大语言模型在整体表现上领先,尤其在生成类任务中;少样本情形下,其性能接近全量数据表现,而小型编码器模型在简单任务中依然有竞争力。跨语言训练多个非目标语言对微调大模型的迁移效果最强,而小型编码器或序列到序列模型则更受益于代码混写,凸显了不同架构适配的跨语言策略。此外,本文还构建了两个新的德语数据集:经调整的GERestaurant和首个德语ASQP数据集GERest,以推动非英语环境下的多语言ABSA研究。
原文摘要 · Abstract (English)
Aspect-based Sentiment Analysis (ABSA) extracts fine-grained opinions toward specific aspects within text but remains largely English-focused despite major advances in transformer-based and instruction-tuned models. This work presents a multilingual evaluation of state-of-the-art ABSA approaches across seven languages (English, German, French, Dutch, Russian, Spanish, and Czech) and four subtasks (ACD, ACSA, TASD, ASQP). We systematically compare different transformer architectures under zero-resource, data-only, and full-resource settings, using cross-lingual transfer, code-switching and machine translation. Fine-tuned Large Language Models (LLMs) achieve the highest overall scores, particularly in complex generative tasks, while few-shot counterparts approach this performance in simpler setups, where smaller encoder models also remain competitive. Cross-lingual training on multiple non-target languages yields the strongest transfer for fine-tuned LLMs, while smaller encoder or seq-to-seq models benefit most from code-switching, highlighting architecture-specific strategies for multilingual ABSA. We further contribute two new German datasets, an adapted GERestaurant and the first German ASQP dataset (GERest), to encourage multilingual ABSA research beyond English.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。