arXiv:2410.11526cs.HCcs.CL2024-10被引 1

人机协作构建粤语情感词典,提升低资源语言情感分析效果。

Human-LLM Collaborative Construction of a Cantonese Emotion Lexicon

  • 结合LLM与人工标注,融合多源语料构建粤语情感词典。
  • 在三个情感数据集上验证词典一致性,提升情感识别准确率。
  • 适合低资源语言研究者与跨语言情感分析从业者参考。

大型语言模型(LLMs)在语言理解与生成方面展现出卓越能力,其内部知识被广泛用于自动化标注。本研究提出通过人机协作方式构建粤语情感词典,针对低资源语言挑战。利用LLM提供的情感标签及人工标注,整合其他语言的词典与本地论坛语料,构建包含口语表达的粤语情感词典。通过修改并使用三个不同的情感文本数据集,评估该词典在情感提取任务中的一致性。结果验证了词典的有效性,强调人机协同标注可显著提升情感标签质量,凸显该合作模式在低资源语言自然语言处理中的潜力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in language understanding and generation. Advanced utilization of the knowledge embedded in LLMs for automated annotation has consistently been explored. This study proposed to develop an emotion lexicon for Cantonese, a low-resource language, through collaborative efforts between LLM and human annotators. By integrating emotion labels provided by LLM and human annotators, the study leveraged existing linguistic resources including lexicons in other languages and local forums to construct a Cantonese emotion lexicon enriched with colloquial expressions. The consistency of the proposed emotion lexicon in emotion extraction was assessed through modification and utilization of three distinct emotion text datasets. This study not only validates the efficacy of the constructed lexicon but also emphasizes that collaborative annotation between human and artificial intelligence can significantly enhance the quality of emotion labels, highlighting the potential of such partnerships in facilitating natural language processing tasks for low-resource languages.

情感分析粤语人机协作低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。