arXiv:2505.16386cs.LG2025-05被引 1

用全状态空间提升可解释性,实现高性能文本嵌入

Omni TM-AE: A Scalable and Interpretable Embedding Model Using the Full Tsetlin Machine State Space

  • 利用Tsetlin机完整状态矩阵,包含以往忽略的字面量
  • 单次训练即可生成可复用的可解释嵌入表示
  • 在语义相似、情感分类等任务中表现优于主流模型

大型语言模型日益复杂,引发对其可解释性和可复用性的担忧。传统嵌入模型如Word2Vec和GloVe虽具可扩展性,但缺乏透明度,常被视为黑箱。相比之下,可解释模型如Tsetlin Machine(TM)虽有潜力构建可解释学习系统,但此前在可扩展性和可复用性上受限。本文提出Omni Tsetlin Machine AutoEncoder(Omni TM-AE),一种新型嵌入模型,充分挖掘TM状态矩阵中的全部信息,包括以往未用于子句构建的字面量。该方法通过一次训练即可构建可复用且可解释的嵌入表示。在语义相似性、情感分类和文档聚类等多项任务上的大量实验表明,Omni TM-AE在性能上与主流嵌入模型相当甚至更优。结果表明,在现代自然语言处理系统中,无需依赖不透明架构,即可实现性能、可扩展性与可解释性的平衡。

原文摘要 · Abstract (English)

The increasing complexity of large-scale language models has amplified concerns regarding their interpretability and reusability. While traditional embedding models like Word2Vec and GloVe offer scalability, they lack transparency and often behave as black boxes. Conversely, interpretable models such as the Tsetlin Machine (TM) have shown promise in constructing explainable learning systems, though they previously faced limitations in scalability and reusability. In this paper, we introduce Omni Tsetlin Machine AutoEncoder (Omni TM-AE), a novel embedding model that fully exploits the information contained in the TM's state matrix, including literals previously excluded from clause formation. This method enables the construction of reusable, interpretable embeddings through a single training phase. Extensive experiments across semantic similarity, sentiment classification, and document clustering tasks show that Omni TM-AE performs competitively with and often surpasses mainstream embedding models. These results demonstrate that it is possible to balance performance, scalability, and interpretability in modern Natural Language Processing (NLP) systems without resorting to opaque architectures.

可解释性嵌入模型Tsetlin机NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。