构建首个带观点词标注的捷克语餐厅评论数据集,推动低资源语言情感分析研究
Extending Czech Aspect-Based Sentiment Analysis with Opinion Terms: Dataset and LLM Benchmarks
- 构建包含观点词标注的捷克语餐厅评论数据集,支持三种ABSA任务
- 在多语言设置下测试模型,发现大模型在处理捷克语细微情感时表现受限
- 提出基于LLM的翻译与标签对齐方法,可推广至其他低资源语言
本文提出首个面向捷克语餐厅领域的方面级情感分析(ABSA)数据集,该数据集通过标注观点词支持三种不同复杂度的ABSA任务。基于此数据集,我们在单语、跨语言及多语言场景下对现代Transformer模型(包括大语言模型)进行了广泛实验。为应对跨语言挑战,我们提出一种利用大语言模型进行翻译与标签对齐的方法,显著提升性能。实验结果揭示了当前先进模型在处理捷克语等低资源语言时的优劣,尤其是对微妙观点词和复杂情感表达的识别能力不足。详细错误分析指出了关键难点。本数据集为捷克语ABSA建立了新基准,所提方法也为其他低资源语言的ABSA资源适配提供可扩展解决方案。
原文摘要 · Abstract (English)
This paper introduces a novel Czech dataset in the restaurant domain for aspect-based sentiment analysis (ABSA), enriched with annotations of opinion terms. The dataset supports three distinct ABSA tasks involving opinion terms, accommodating varying levels of complexity. Leveraging this dataset, we conduct extensive experiments using modern Transformer-based models, including large language models (LLMs), in monolingual, cross-lingual, and multilingual settings. To address cross-lingual challenges, we propose a translation and label alignment methodology leveraging LLMs, which yields consistent improvements. Our results highlight the strengths and limitations of state-of-the-art models, especially when handling the linguistic intricacies of low-resource languages like Czech. A detailed error analysis reveals key challenges, including the detection of subtle opinion terms and nuanced sentiment expressions. The dataset establishes a new benchmark for Czech ABSA, and our proposed translation-alignment approach offers a scalable solution for adapting ABSA resources to other low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。