用可配置框架让形式概念自动取名,提升可读性
A Variability-Based Framework for Interpretable Naming in Formal and Relational Concept Analysis

- 基于变异性模型控制命名信息源,灵活配置上下文
- 在披萨店数据集上验证,不同配置生成不同语义名称
- 帮助专家理解概念关系,发现建模潜在问题
从符号数据中提取知识常产生形式定义明确但难以理解的抽象。形式概念分析(FCA)与关系概念分析(RCA)能从对象描述和关系中生成显式的概念结构、蕴含关系与关联依赖,虽具可解释性,但其概念多以技术标签标识,限制了作为人类可理解知识单元的应用。为解决这一问题,本文从符号知识表示视角研究FCA与RCA中的概念命名。首先识别命名面临的语言与术语挑战,包括歧义、区分度、简洁性及概念间一致性。随后提出一种可配置的LLM辅助命名框架,通过变异性模型控制命名时暴露的信息源,如意图、扩展、继承信息、邻近概念、蕴含关系与关系属性,从而显式揭示从形式描述到人可读名称的语义选择过程。以小规模披萨店关系数据集为例进行演示,结果表明不同配置影响LLM生成名称,命名变异性可揭示解释选择、关系依赖及底层符号数据中的建模缺陷。
原文摘要 · Abstract (English)
Knowledge extraction from symbolic data often produces abstractions that are formally defined but not immediately interpretable by users. Formal Concept Analysis (FCA) and Relational Concept Analysis (RCA) provide representative settings for this issue: they generate explicit conceptual structures, implications, and relational dependencies from object descriptions and relations. Although these structures are explainable by design, their concepts are often identified by technical labels, which limits their use as human-interpretable knowledge units. Assigning meaningful names to such concepts is therefore a key issue for interpretation, navigation, validation, and reuse by domain experts. This paper investigates concept naming in FCA and RCA from a symbolic knowledge representation perspective. We first characterize the linguistic and terminological challenges involved in naming generated symbolic abstractions, including ambiguity, discrimination, concision, and consistency across related concepts. We then propose a configurable framework for LLM-assisted concept naming. The framework relies on a variability model that controls which sources of information are exposed during naming, such as intent, extent, inherited information, neighboring concepts, implications, and relational attributes. It thereby makes explicit the semantic choices involved in moving from formal concept descriptions to human-readable names. The approach is illustrated as a proof of concept on a small relational dataset in the pizzeria domain. This illustration shows how different configurations influence the names suggested by an LLM, and how naming variability can reveal interpretation choices, relational dependencies, and possible modeling issues in the underlying symbolic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。