测试大模型在波斯语情绪识别中自解释的可信度
Can LLMs Faithfully Explain Themselves in Low-Resource Languages? A Case Study on Emotion Detection in Persian
- 用词重要性对比法评估模型解释与人工判断的一致性
- 模型分类准确但解释与人类判断差异大,自洽性强于真实性
- 提示顺序影响解释质量,现有方法在低资源语言中不可靠
大语言模型(LLMs)越来越常生成预测的同时附带自解释,但这些解释的可信度在低资源语言中令人担忧。本研究以波斯语情绪分类为例,通过比较模型识别的关键词与人工标注者的结果,评估解释的可信度。采用基于标记级对数概率的置信度分数作为评估指标。测试了两种提示策略:先预测再解释(Predict-then-Explain)和先解释再预测(Explain-then-Predict),分析其对解释可信度的影响。结果表明,尽管模型分类性能强,其生成的解释往往偏离真实推理,与自身一致性高于与人类判断的一致性。这揭示了当前解释方法与评估指标的局限性,强调需发展更可靠的多语言、低资源场景下的模型可靠性机制。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to generate self-explanations alongside their predictions, a practice that raises concerns about the faithfulness of these explanations, especially in low-resource languages. This study evaluates the faithfulness of LLM-generated explanations in the context of emotion classification in Persian, a low-resource language, by comparing the influential words identified by the model against those identified by human annotators. We assess faithfulness using confidence scores derived from token-level log-probabilities. Two prompting strategies, differing in the order of explanation and prediction (Predict-then-Explain and Explain-then-Predict), are tested for their impact on explanation faithfulness. Our results reveal that while LLMs achieve strong classification performance, their generated explanations often diverge from faithful reasoning, showing greater agreement with each other than with human judgments. These results highlight the limitations of current explanation methods and metrics, emphasizing the need for more robust approaches to ensure LLM reliability in multilingual and low-resource contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。