无需标注自动发现可理解的文本解释概念。
Enhancing the Comprehensibility of Text Explanations via Unsupervised Concept Discovery
- 用对象中心架构自动提取语义概念。
- 大模型评估概念可理解性并指导优化。
- 适合需要可信解释的自然语言应用。
基于概念的可解释AI方法因能契合人类推理而受到关注,但在文本领域应用有限。现有方法多依赖预定义概念标注,无法发现新概念;无监督方法虽能自动提取概念,但解释常难被人类直观理解,削弱用户信任。为此,我们提出ECO-Concept框架,无需任何概念标注即可自动发现可理解的概念。该框架首先利用对象中心架构自动提取语义概念,再通过大语言模型评估概念的可理解性,并依据评估结果指导模型微调,从而获得更易懂的解释。实验表明,本方法在多种任务中表现优异;进一步概念评估显示,其学习到的概念在可理解性上优于当前主流方法。
原文摘要 · Abstract (English)
Concept-based explainable approaches have emerged as a promising method in explainable AI because they can interpret models in a way that aligns with human reasoning. However, their adaption in the text domain remains limited. Most existing methods rely on predefined concept annotations and cannot discover unseen concepts, while other methods that extract concepts without supervision often produce explanations that are not intuitively comprehensible to humans, potentially diminishing user trust. These methods fall short of discovering comprehensible concepts automatically. To address this issue, we propose \textbf{ECO-Concept}, an intrinsically interpretable framework to discover comprehensible concepts with no concept annotations. ECO-Concept first utilizes an object-centric architecture to extract semantic concepts automatically. Then the comprehensibility of the extracted concepts is evaluated by large language models. Finally, the evaluation result guides the subsequent model fine-tuning to obtain more understandable explanations. Experiments show that our method achieves superior performance across diverse tasks. Further concept evaluations validate that the concepts learned by ECO-Concept surpassed current counterparts in comprehensibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。