arXiv:2507.02922cs.LGcs.HC2025-07被引 7

用概念模型提升机器学习准确率与可解释性

Domain Knowledge in Artificial Intelligence: Using Conceptual Modeling to Increase Machine Learning Accuracy and Explainability

  • 基于概念模型构建数据准备指南,融合领域知识
  • 在两个真实场景中验证,模型性能显著提升
  • 适合需要高透明度的医疗、金融等专业领域

机器学习能从海量异构数据中提取有用信息,但其性能与透明性仍存挑战,部分源于对领域知识利用不足。本文提出一种名为概念模型赋能机器学习(CMML)的方法,通过概念建模的结构与原则指导数据预处理。研究在两个真实问题上应用该方法,评估其对模型性能的影响,并征求数据科学家对其适用性的评价。结果表明,CMML可有效提升机器学习效果,增强模型可解释性。

原文摘要 · Abstract (English)

Machine learning enables the extraction of useful information from large, diverse datasets. However, despite many successful applications, machine learning continues to suffer from performance and transparency issues. These challenges can be partially attributed to the limited use of domain knowledge by machine learning models. This research proposes using the domain knowledge represented in conceptual models to improve the preparation of the data used to train machine learning models. We develop and demonstrate a method, called the Conceptual Modeling for Machine Learning (CMML), which is comprised of guidelines for data preparation in machine learning and based on conceptual modeling constructs and principles. To assess the impact of CMML on machine learning outcomes, we first applied it to two real-world problems to evaluate its impact on model performance. We then solicited an assessment by data scientists on the applicability of the method. These results demonstrate the value of CMML for improving machine learning outcomes.

机器学习领域知识可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。