arXiv:2506.23524cs.CLcs.AI2025-06被引 1

构建首个面向越南语教育情感与话题分类的多任务数据集

NEU-ESC: A Comprehensive Vietnamese dataset for Educational Sentiment analysis and topic Classification toward multitask learning

  • 从大学论坛收集数据,覆盖更长文本和丰富词汇
  • 基于BERT的多任务学习达83.7%情感分类准确率
  • 适合研究越南语NLP及教育文本分析的学者使用

在教育领域,通过学生评论理解其意见至关重要,尤其在越南语中资源仍较匮乏。现有教育数据集常缺乏领域相关性与学生俚语。为此,我们提出NEU-ESC,一个源自大学论坛的越南语教育情感分类与话题分类数据集,样本更多、类别更丰富、文本更长、词汇更广。我们还探索了基于编码器的预训练模型(如BERT)的多任务学习,结果显示情感分类与话题分类准确率分别达到83.7%和79.8%。同时,我们对数据集与模型在多个基准上进行了评估,包括大语言模型,并展开讨论。该数据集已公开:https://huggingface.co/datasets/hung20gg/NEU-ESC。

原文摘要 · Abstract (English)

In the field of education, understanding students' opinions through their comments is crucial, especially in the Vietnamese language, where resources remain limited. Existing educational datasets often lack domain relevance and student slang. To address these gaps, we introduce NEU-ESC, a new Vietnamese dataset for Educational Sentiment Classification and Topic Classification, curated from university forums, which offers more samples, richer class diversity, longer texts, and broader vocabulary. In addition, we explore multitask learning using encoder-only language models (BERT), in which we showed that it achieves performance up to 83.7% and 79.8% accuracy for sentiment and topic classification tasks. We also benchmark our dataset and model with other datasets and models, including Large Language Models, and discuss these benchmarks. The dataset is publicly available at: https://huggingface.co/datasets/hung20gg/NEU-ESC.

情感分析越南语多任务学习教育NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。