arXiv:2608.12776cs.CL2026-08中稿 · publication at 202…

构建越南社交文本目标情感检测数据集,助力本地化情感分析研究

ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts

  • 构建10,985条评论的标注数据集,含21,244个情感四元组
  • 发现越南语隐含主体与目标、词汇歧义等特有现象
  • 适合从事越南语情感分析、多语言NLP的研究者参考

本文提出面向越南社交媒体文本的目标导向情感检测数据集ViTOED。该数据集包含10,985条用户评论和21,244个手工标注的情感四元组(来源、目标、表达、极性),遵循严格标注规范。数据揭示了越南语中隐含来源与目标、词汇歧义等独特现象,有助于深入分析用户对实体的情感。我们提出基于结构化情感图的基线方法,并评估多种越南预训练语言模型。实证结果表明,在跨度检测与关系抽取任务上仍存在显著挑战,模型在越南目标导向情感检测任务中仍有较大提升空间。

原文摘要 · Abstract (English)

This paper introduces ViTOED, a novel dataset for target-oriented emotion detection in Vietnamese social media texts. The ViTOED comprises 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity) that follow strict guidelines. The dataset reveals Vietnamese-specific phenomena, such as implicit sources and targets and vocabulary ambiguities, enabling deeper analysis of user emotions toward entities. We propose a baseline using structured sentiment graphs and evaluate various Vietnamese pre-trained language models. The empirical results highlight challenges in span detection and relation extraction and indicate substantial room for model improvement in Vietnamese Target-Oriented Emotion Detection tasks.

情感分析越南语数据集NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。