arXiv:2410.03296cs.CLcs.AI2024-10中稿 · ACL被引 1

对比大模型自解释与人工理由,发现解释质量受文本长度和任务复杂度影响。

A Systematic Comparison between Extractive Self-Explanations and Human Rationales in Text Classification

  • 用输入级理由评估大模型自解释的合理性,对比人类标注。
  • 自解释在长文本和复杂任务中与人工理由一致性下降。
  • 自解释聚焦关键语义词,后处理方法则偏向格式化符号。

指令微调的大语言模型可生成自解释以说明其输出,无需复杂可解释性技术。本文分析该能力是否产生高质量解释,通过输入级理由评估其对人类的合理性。研究涵盖情感分类、强迫劳动检测和主张验证三类文本分类任务,包含丹麦语和意大利语的情感分类数据,并收集了Climate-Fever数据集的人工理由标注。进一步评估人类与自解释理由在正确预测下的忠实性,并引入后处理归因解释进行对比。分析四个开源大模型发现,自解释与人类理由的一致性高度依赖文本长度和任务复杂度。尽管如此,自解释仍能生成忠实的词级别理由子集;而后处理归因方法则倾向于强调结构和格式化标记,体现根本不同的解释策略。

原文摘要 · Abstract (English)

Instruction-tuned LLMs are able to provide \textit{an} explanation about their output to users by generating self-explanations, without requiring the application of complex interpretability techniques. In this paper, we analyse whether this ability results in a \textit{good} explanation. We evaluate self-explanations in the form of input rationales with respect to their plausibility to humans. We study three text classification tasks: sentiment classification, forced labour detection and claim verification. We include Danish and Italian translations of the sentiment classification task and compare self-explanations to human annotations. For this, we collected human rationale annotations for Climate-Fever, a claim verification dataset. We furthermore evaluate the faithfulness of human and self-explanation rationales with respect to correct model predictions, and extend the study by incorporating post-hoc attribution-based explanations. We analyse four open-weight LLMs and find that alignment between self-explanations and human rationales highly depends on text length and task complexity. Nevertheless, self-explanations yield faithful subsets of token-level rationales, whereas post-hoc attribution methods tend to emphasize structural and formatting tokens, reflecting fundamentally different explanation strategies.

大模型解释自解释可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。