arXiv:2604.07057cs.CL2026-04被引 2

让印尼语情感分析模型理解话题背景,提升判断准确性。

IndoBERT-Sentiment: Context-Conditioned Sentiment Classification for Indonesian Text

  • 输入话题上下文和文本,联合判断情感
  • 在188个话题上达到F1宏值0.856,准确率88.1%
  • 对传统方法误判的句子显著改善,适合多话题场景

现有印尼语情感分析模型孤立处理文本,忽略常决定情感倾向的话题背景。我们提出IndoBERT-Sentiment,一种基于话题上下文的情感分类器,同时接收话题上下文与待分析文本,输出与话题相关的合理情感判断。该模型基于IndoBERT Large(335M参数),在包含31,360组上下文-文本对、覆盖188个话题的数据集上训练,取得F1宏值0.856与准确率88.1%。在与三种主流通用印尼语情感模型的同测试集对比中,其表现领先最佳基线达35.6个F1点。实验表明,此前仅用于相关性分类的话题条件化机制,可有效迁移至情感分析任务,使模型正确识别出传统无上下文方法系统误判的文本。

原文摘要 · Abstract (English)

Existing Indonesian sentiment analysis models classify text in isolation, ignoring the topical context that often determines whether a statement is positive, negative, or neutral. We introduce IndoBERT-Sentiment, a context-conditioned sentiment classifier that takes both a topical context and a text as input, producing sentiment predictions grounded in the topic being discussed. Built on IndoBERT Large (335M parameters) and trained on 31,360 context-text pairs labeled across 188 topics, the model achieves an F1 macro of 0.856 and accuracy of 88.1%. In a head-to-head evaluation against three widely used general-purpose Indonesian sentiment models on the same test set, IndoBERT-Sentiment outperforms the best baseline by 35.6 F1 points. We show that context-conditioning, previously demonstrated for relevancy classification, transfers effectively to sentiment analysis and enables the model to correctly classify texts that are systematically misclassified by context-free approaches.

情感分析上下文建模印尼语IndoBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。