对比大模型自解释与人工理由,发现解释质量受文本长度和任务复杂度影响。
A Systematic Comparison between Extractive Self-Explanations and Human Rationales in Text Classification
- 用输入级理由评估大模型自解释的合理性,对比人类标注。
- 自解释在长文本和复杂任务中与人工理由一致性下降。
- 自解释聚焦关键语义词,后处理方法则偏向格式化符号。
指令微调的大语言模型可生成自解释以说明其输出,无需复杂可解释性技术。本文分析该能力是否产生高质量解释,通过输入级理由评估其对人类的合理性。研究涵盖情感分类、强迫劳动检测和主张验证三类文本分类任务,包含丹麦语和意大利语的情感分类数据,并收集了Climate-Fever数据集的人工理由标注。进一步评估人类与自解释理由在正确预测下的忠实性,并引入后处理归因解释进行对比。分析四个开源大模型发现,自解释与人类理由的一致性高度依赖文本长度和任务复杂度。尽管如此,自解释仍能生成忠实的词级别理由子集;而后处理归因方法则倾向于强调结构和格式化标记,体现根本不同的解释策略。
原文摘要 · Abstract (English)
Instruction-tuned LLMs are able to provide \textit{an} explanation about their output to users by generating self-explanations, without requiring the application of complex interpretability techniques. In this paper, we analyse whether this ability results in a \textit{good} explanation. We evaluate self-explanations in the form of input rationales with respect to their plausibility to humans. We study three text classification tasks: sentiment classification, forced labour detection and claim verification. We include Danish and Italian translations of the sentiment classification task and compare self-explanations to human annotations. For this, we collected human rationale annotations for Climate-Fever, a claim verification dataset. We furthermore evaluate the faithfulness of human and self-explanation rationales with respect to correct model predictions, and extend the study by incorporating post-hoc attribution-based explanations. We analyse four open-weight LLMs and find that alignment between self-explanations and human rationales highly depends on text length and task complexity. Nevertheless, self-explanations yield faithful subsets of token-level rationales, whereas post-hoc attribution methods tend to emphasize structural and formatting tokens, reflecting fundamentally different explanation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。