arXiv:2505.09286cs.CL2025-05综述被引 3

无需标注数据,自动为多语言多领域评论打标签

A Scalable Unsupervised Framework for multi-aspect labeling of Multilingual and Multi-Domain Review Data

  • 通过聚类与负采样生成带方面感知的嵌入向量
  • 自动生成标签在多个语言和领域上表现优异
  • 适合需要大规模无监督评论分析的研究者

有效分析在线评论对各行业至关重要,但现有研究多局限于特定语言或领域,且依赖需大量标注数据的有监督方法。为此,我们提出一种可扩展的多语言、无监督框架,用于跨领域方面检测。该框架对韩语和英语多领域评论数据进行自动标注,并通过大量实验评估生成标签的质量。首先通过聚类提取方面类别候选,再利用负采样将每篇评论表示为方面感知的嵌入向量。为评估效果,我们在多方面标注任务上微调多个预训练语言模型,结果表明模型性能优异,证明自动生成标签具备训练价值。与公开的大语言模型对比显示,本框架在处理大规模数据时具更高一致性与可扩展性。人工评估也证实自动标签质量接近人工标注水平。本研究展示了克服有监督方法局限、适应多语言多领域环境的强健方面标注方案潜力。未来工作将探索自动评论摘要及人工智能代理集成,以提升分析效率与深度。

原文摘要 · Abstract (English)

Effectively analyzing online review data is essential across industries. However, many existing studies are limited to specific domains and languages or depend on supervised learning approaches that require large-scale labeled datasets. To address these limitations, we propose a multilingual, scalable, and unsupervised framework for cross-domain aspect detection. This framework is designed for multi-aspect labeling of multilingual and multi-domain review data. In this study, we apply automatic labeling to Korean and English review datasets spanning various domains and assess the quality of the generated labels through extensive experiments. Aspect category candidates are first extracted through clustering, and each review is then represented as an aspect-aware embedding vector using negative sampling. To evaluate the framework, we conduct multi-aspect labeling and fine-tune several pretrained language models to measure the effectiveness of the automatically generated labels. Results show that these models achieve high performance, demonstrating that the labels are suitable for training. Furthermore, comparisons with publicly available large language models highlight the framework's superior consistency and scalability when processing large-scale data. A human evaluation also confirms that the quality of the automatic labels is comparable to those created manually. This study demonstrates the potential of a robust multi-aspect labeling approach that overcomes limitations of supervised methods and is adaptable to multilingual, multi-domain environments. Future research will explore automatic review summarization and the integration of artificial intelligence agents to further improve the efficiency and depth of review analysis.

无监督学习多语言评论分析方面检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。