arXiv:2505.19254cs.CL2025-05ACL

用NLP分析用户评论,识别同一产品在不同市场的质量差异

Unveiling Dual Quality in Product Reviews: An NLP-Based Approach

  • 构建波兰语评论数据集,含1957条带双质量问题的评论
  • 基于SetFit与LLM的方法在识别质量差异上表现优异
  • 支持多语言迁移,适合跨境电商平台质量监控

消费者常面临产品品质不一致的问题,尤其是同一产品在不同市场存在差异,即所谓的双质量问题。为识别并应对该问题,亟需自动化技术。本文探讨自然语言处理(NLP)在检测此类差异中的应用,并完整呈现解决方案的开发流程。首先,我们详细介绍了新构建的波兰语数据集,包含1,957条评论,其中540条明确涉及双质量问题。随后,我们对多种方法进行实验,包括SetFit结合sentence-transformers、基于Transformer的编码器以及大语言模型(LLMs),并开展误差分析与鲁棒性验证。此外,还评估了在英语、法语和德语子集上的多语言迁移效果。最后,论文总结了部署建议与实际应用场景。

原文摘要 · Abstract (English)

Consumers often face inconsistent product quality, particularly when identical products vary between markets, a situation known as the dual quality problem. To identify and address this issue, automated techniques are needed. This paper explores how natural language processing (NLP) can aid in detecting such discrepancies and presents the full process of developing a solution. First, we describe in detail the creation of a new Polish-language dataset with 1,957 reviews, 540 highlighting dual quality issues. We then discuss experiments with various approaches like SetFit with sentence-transformers, transformer-based encoders, and LLMs, including error analysis and robustness verification. Additionally, we evaluate multilingual transfer using a subset of opinions in English, French, and German. The paper concludes with insights on deployment and practical applications.

NLP评论分析双质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。