用数据驱动方法整合信息系统中概念及其关系,提升知识累积性
GUT-IS: A Data-Driven Approach to Integrating Constructs and Their Relations in Information Systems

- 结合文本嵌入与聚类生成概念分组候选方案
- 通过权衡语义纯度与聚类数量,选出最优分组结果
- 适用于需要统一概念定义的管理学与信息系统研究
结构方程模型在信息系统研究中广泛应用,但概念定义不一致阻碍了知识的累积发展。本文提出一种将结构方程模型整合为统一模型的数据驱动方法:利用任务适配的文本嵌入与聚类生成概念分组候选集;随后通过显式权衡语义纯度与聚类数量的损失函数选择最优解。该方法可分析当优先级从语义纯度转向聚类简洁性时,概念分组及其关系的变化情况。我们在两个信息系统领域数据集上对所提方法进行了实证评估与探索。
原文摘要 · Abstract (English)
Structural equation modeling is widely used in IS research. However, inconsistent construct definitions impede the cumulative development of knowledge. In this work, we present an approach that aims at the integration of structural equation models into a unified model: We use a combination of task-adapted text embeddings and clustering to produce a candidate set of construct groupings. Subsequently, we select the optimal solution using a loss function that explicitly trades off semantic purity and parsimony in the number of clusters. By making this trade-off explicit, our approach allows to analyze how construct groupings and their relations change as one shifts the priority from purity to parsimony. Empirically, we evaluate and explore the proposed methodology on two datasets from the IS domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。