arXiv:2511.16353cs.CLcs.AI2025-11中稿 · the main conferenc…被引 1

探究可解释性与模型性能的关系,发现高信息量的解释未必提升准确率。

Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies

  • 通过词级分类和注意力正则化分析解释的可靠性
  • 高充分性解释反而降低分类准确率,因上下文干扰
  • 跨领域分类中解释信息可提升性能,但效果不一致

人类对自然语言的解释(即理由)可用于评估模型是否基于正确原因学习标签,而非依赖数据集特定捷径。充分性是衡量理由信息量的常用指标,但无法揭示理由信息对模型性能的影响。本文将充分性与两种建模范式关联:模型识别哪些词属于理由(词级分类)以及通过注意力正则化在输入中融入理由以提升性能。结果表明,高信息量的理由并不利于正确分类;充分性实际反映的是非理由上下文对分类的影响,而该上下文会干扰理由信息。此外,在输入中引入理由信息可提升跨领域分类性能,但结果因任务和模型类型而异。最后,充分性与词级分类能力无明显关联。这些发现揭示了理由的复杂性,提示需发展更系统的方法来捕捉此类信息。

原文摘要 · Abstract (English)

Human explanations of natural language, rationales, form a tool to assess whether models learn a label for the right reasons or rely on dataset-specific shortcuts. Sufficiency is a common metric for estimating the informativeness of rationales, but it provides limited insight into the effects of rationale information on model performance. We address this limitation by relating sufficiency to two modelling paradigms: the ability of models to identify which tokens are part of the rationale (through token classification) and the ability of improving model performance by incorporating rationales in the input (through attention regularisation). We find that highly informative rationales are not likely to help classify the instance correctly. Sufficiency conversely captures the classification impact of the non-rationalised context, which interferes with rationale information in the same input. We also find that incorporating rationale information in model inputs can boost cross-domain classification, but results are inconsistent per task and model type. Finally, sufficiency and token classification appear to be unrelated. These results exemplify the complexity of rationales, showing that metrics capable of systematically capturing this type of information merit further investigation.

可解释性模型评估自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。