用结构化知识提升语言模型生成结果的可解释性
Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
- 通过实体与知识三元组两级结构设计通用知识搜索器
- 在两个数据集上优于现有方法,生成过程更可解释
- 适合需要可靠、可解释生成结果的应用场景
知识增强型文本生成旨在利用内部或外部知识源提升生成文本质量。尽管语言模型在生成连贯流畅文本方面表现优异,但其缺乏可解释性成为重大障碍,尤其在要求高可靠性与可解释性的知识增强生成任务中。现有方法多依赖特定领域知识检索器,难以泛化到不同数据类型与任务。为此,我们直接利用结构化知识的两层架构(高层实体与低层知识三元组),设计了任务无关的结构化知识猎手。具体采用局部-全局交互机制进行知识表征学习,并以分层Transformer指针网络作为核心,选择相关知识三元组与实体。结合语言模型强生成能力与知识猎手的高忠实度,本模型实现高可解释性,使用户能理解生成过程。实验证明,该模型在RotoWireFG数据集上的内部知识表格转文本任务与KdConv数据集上的外部知识对话生成任务中均显著优于当前最优方法及基线语言模型,在基准测试上树立新标准。
原文摘要 · Abstract (English)
Knowledge-enhanced text generation aims to enhance the quality of generated text by utilizing internal or external knowledge sources. While language models have demonstrated impressive capabilities in generating coherent and fluent text, the lack of interpretability presents a substantial obstacle. The limited interpretability of generated text significantly impacts its practical usability, particularly in knowledge-enhanced text generation tasks that necessitate reliability and explainability. Existing methods often employ domain-specific knowledge retrievers that are tailored to specific data characteristics, limiting their generalizability to diverse data types and tasks. To overcome this limitation, we directly leverage the two-tier architecture of structured knowledge, consisting of high-level entities and low-level knowledge triples, to design our task-agnostic structured knowledge hunter. Specifically, we employ a local-global interaction scheme for structured knowledge representation learning and a hierarchical transformer-based pointer network as the backbone for selecting relevant knowledge triples and entities. By combining the strong generative ability of language models with the high faithfulness of the knowledge hunter, our model achieves high interpretability, enabling users to comprehend the model output generation process. Furthermore, we empirically demonstrate the effectiveness of our model in both internal knowledge-enhanced table-to-text generation on the RotoWireFG dataset and external knowledge-enhanced dialogue response generation on the KdConv dataset. Our task-agnostic model outperforms state-of-the-art methods and corresponding language models, setting new standards on the benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。