针对菲律宾电商评论中的混语现象,提出一套高效的情感方面抽取方案。
Aspect Extraction from E-Commerce Product and Service Reviews
- 融合规则、大模型和微调模型的多方法流水线,处理显式与隐式方面。
- 生成式大模型在所有任务中表现最佳,宏平均F1达0.91。
- 适用于低资源及混语环境,对跨语言电商分析有实用价值。
方面抽取(AE)是基于方面的情感分析(ABSA)的关键任务,但在低资源和混语场景(如菲律宾常见的塔加洛语与英语混合的Taglish)中仍具挑战。本文提出一个专为Taglish设计的完整AE流程,结合规则、大语言模型(LLM)和微调技术,实现方面识别与提取。通过多方法主题建模构建分层方面框架(HAF),并采用双模式标注机制处理显式与隐式方面。评估了四种模型:规则系统、生成式大模型(Gemini 2.0 Flash)以及两个在不同数据集上微调的Gemma-3 1B模型(分别基于规则与大模型标注)。结果表明,生成式大模型在所有任务中表现最优,宏平均F1达0.91,尤其擅长处理隐式方面;而微调模型受限于数据分布不均与模型容量,表现有限。本研究贡献了一个可扩展、语言自适应的框架,助力多样化混语环境中ABSA的提升。
原文摘要 · Abstract (English)
Aspect Extraction (AE) is a key task in Aspect-Based Sentiment Analysis (ABSA), yet it remains difficult to apply in low-resource and code-switched contexts like Taglish, a mix of Tagalog and English commonly used in Filipino e-commerce reviews. This paper introduces a comprehensive AE pipeline designed for Taglish, combining rule-based, large language model (LLM)-based, and fine-tuning techniques to address both aspect identification and extraction. A Hierarchical Aspect Framework (HAF) is developed through multi-method topic modeling, along with a dual-mode tagging scheme for explicit and implicit aspects. For aspect identification, four distinct models are evaluated: a Rule-Based system, a Generative LLM (Gemini 2.0 Flash), and two Fine-Tuned Gemma-3 1B models trained on different datasets (Rule-Based vs. LLM-Annotated). Results indicate that the Generative LLM achieved the highest performance across all tasks (Macro F1 0.91), demonstrating superior capability in handling implicit aspects. In contrast, the fine-tuned models exhibited limited performance due to dataset imbalance and architectural capacity constraints. This work contributes a scalable and linguistically adaptive framework for enhancing ABSA in diverse, code-switched environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。