arXiv:2509.21241cs.LGcs.AI2025-09被引 1

用知识图谱生成反事实,解释微调后的大模型如何改变推理结构。

Explaining Fine Tuned LLMs via Counterfactuals A Knowledge Graph Driven Framework

  • 基于知识图谱构建反事实扰动,生成最小结构变化
  • 揭示微调后模型的结构依赖与参数偏移的一致性
  • 适合关注大模型可解释性的研究人员

低秩适配(LoRA)的广泛应用使大语言模型(LLMs)能高效获取领域知识,但其微调机制如何影响模型的结构推理与语义行为仍不明确。本文提出一种基于知识图谱的反事实解释框架,构建生物信息学工具领域的异质知识图谱BioToolKG,设计CFFTLLMExplainer方法,通过学习图节点与边上的软掩码,生成最小结构扰动以引发最大语义差异。该方法联合优化结构稀疏性与语义偏离,并引入熵正则化与边平滑约束以保持可解释性。在基于LLaMA微调的模型上应用,结果表明反事实掩码能暴露模型结构依赖,且与LoRA引起的参数偏移一致。本工作为理解微调后大模型内部机制提供了新视角,凸显反事实图谱在可解释人工智能中的潜力。

原文摘要 · Abstract (English)

The widespread adoption of Low-Rank Adaptation (LoRA) has enabled large language models (LLMs) to acquire domain-specific knowledge with remarkable efficiency. However, understanding how such a fine-tuning mechanism alters a model's structural reasoning and semantic behavior remains an open challenge. This work introduces a novel framework that explains fine-tuned LLMs via counterfactuals grounded in knowledge graphs. Specifically, we construct BioToolKG, a domain-specific heterogeneous knowledge graph in bioinformatics tools and design a counterfactual-based fine-tuned LLMs explainer (CFFTLLMExplainer) that learns soft masks over graph nodes and edges to generate minimal structural perturbations that induce maximum semantic divergence. Our method jointly optimizes structural sparsity and semantic divergence while enforcing interpretability preserving constraints such as entropy regularization and edge smoothness. We apply this framework to a fine-tuned LLaMA-based LLM and reveal that counterfactual masking exposes the model's structural dependencies and aligns with LoRA-induced parameter shifts. This work provides new insights into the internal mechanisms of fine-tuned LLMs and highlights counterfactual graphs as a potential tool for interpretable AI.

大模型解释知识图谱反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。