arXiv:2603.12227cs.SCcs.LG2026-03被引 1

用模糊规则解析特定领域中CLIP文本嵌入的语义

Interpreting Contrastive Embeddings in Specific Domains with Fuzzy Rules

  • 结合模糊规则与文本处理技术,映射关键特征到CLIP向量空间
  • 在临床报告和影评数据上验证,可解释性强且效果稳定
  • 适合需要可解释性的人工智能应用,如医疗与内容分析

自由文本仍是法律程序和医疗记录等真实场景中数据登记的常见方式。为此,自然语言处理领域致力于将这些文本转化为机器学习可用的结构化格式。目前最流行的文本向量化方法之一是基于图文对比训练的CLIP模型。尽管CLIP在零样本和少样本学习中表现优异,但在特定领域仍存在局限。本文采用模糊规则分类系统,结合标准文本处理技术,将关注特征映射至CLIP生成的向量空间,并分析所得规则与特征重要性。实验在临床报告与影评两个领域展开,分别评估并比较单领域与联合领域的结果。最后讨论该方法的局限性及未来改进方向。

原文摘要 · Abstract (English)

Free-style text is still one of the common ways in which data is registered in real environments, like legal procedures and medical records. Because of that, there have been significant efforts in the area of natural language processing to convert these texts into a structured format, which standard machine learning methods can then exploit. One of the most popular methods to embed text into a vectorial representation is the Contrastive Language-Image Pre-training model (CLIP), which was trained using both image and text. Although the representations computed by CLIP have been very successful in zero-show and few-shot learning problems, they still have problems when applied to a particular domain. In this work, we use a fuzzy rule-based classification system along with some standard text procedure techniques to map some of our features of interest to the space created by a CLIP model. Then, we discuss the rules and associations obtained and the importance of each feature considered. We apply this approach in two different data domains, clinical reports and film reviews, and compare the results obtained individually and when considering both. Finally, we discuss the limitations of this approach and how it could be further improved.

可解释性CLIP模糊逻辑领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。