arXiv:2511.09313cs.CL2025-11

让高棉语情感分类模型能自解释,找出判断依据。

Towards Explainable Khmer Polarity Classification

  • 用指令微调Qwen-3模型,使其生成预测理由。
  • 在新数据集上准确分类并识别关键情感词。
  • 适合需要透明决策的高棉语自然语言应用。

高棉语情感分类是为给定文本分配正向、负向或中性标签的基础自然语言处理任务。现有高棉语模型通常仅输出标签而无法解释推理过程。本文提出一种可解释的高棉语情感分类器,通过微调基于指令的推理模型Qwen-3实现。本文中的可解释性限于模型自身的自解释,即用关键词或短语来说明预测依据。实验表明,微调后的模型不仅能准确预测情感标签,还能识别支持判断的极性相关词汇。此外,本文构建了一个新的高棉语情感数据集,包含短至中等长度的非正式、罗马化及混合编码表达,该数据集通过启发式规则与人工校对创建,可通过受控的Hugging Face仓库(rinabuoy/khmerpolarity_nonreasoning)公开获取。微调后的Qwen-3模型也已同步发布于同一账户。

原文摘要 · Abstract (English)

Khmer polarity classification is a fundamental natural language processing task that assigns a positive, negative, or neutral label to a given Khmer text input. Existing Khmer models typically predict the label without explaining the rationale behind the prediction. This paper proposes an explainable Khmer polarity classifier by fine-tuning an instruction-based reasoning Qwen-3 model. The notion of explainability in this paper is limited to self-explanations, which the model uses to rationalize its predictions. Experimental results show that the fine-tuned model not only predicts labels accurately but also provides reasoning by identifying polarity-related keywords or phrases to support its predictions. In addition, we contribute a new Khmer polarity dataset consisting of short- to medium-length casual, romanized, and mixed-code Khmer expressions. This dataset was constructed using both heuristic rules and human curation and is publicly available through a gated Hugging Face repository (rinabuoy/khmerpolarity_nonreasoning). The fine-tuned Qwen-3 models are also made available in the same Hugging Face account.

情感分析可解释性高棉语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。