arXiv:2502.07039cs.LGcs.HC2025-02被引 3

让人类参与视觉分析,提升分类模型在重叠区域的准确性和可解释性。

Boosting of Classification Models with Human-in-the-Loop Computational Visual Knowledge Discovery

  • 引入人机协同视觉学习,聚焦类别重叠区而非仅错误样本
  • 在无损可视化空间中实现纯区域与重叠区分离分类,准确率达100%
  • 适合医疗等高风险领域,提升用户对模型决策的信任

高风险人工智能分类任务(如医疗诊断)需要高精度且可解释的预测模型。传统分类器为追求整体准确率而牺牲个别案例精度,难以分析类别重叠区域。自适应提升算法虽能通过加权误分样本改进性能,但依赖弱基分类器,限制了性能提升。本文提出将提升方法从仅关注误分类样本扩展至所有类别重叠区域,结合计算与交互式视觉学习(CIVL),利用人类领域知识和视觉洞察构建分类器。采用分而治之的分类流程,将样本分为简单与复杂两类,分别通过计算分析与无损可视化(如平行坐标系)处理。纯区域样本生成可解释的子模型(如命题逻辑或一阶逻辑规则),重叠区域则通过无损可视化降低用户认知负担,识别难判模式并设计新可分类特征。实验显示对鸢尾花数据集可实现100%准确率且完全可解释;模拟数据表明该方法显著提升模型准确率与可解释性,增强用户信心。

原文摘要 · Abstract (English)

High-risk artificial intelligence and machine learning classification tasks, such as healthcare diagnosis, require accurate and interpretable prediction models. However, classifier algorithms typically sacrifice individual case-accuracy for overall model accuracy, limiting analysis of class overlap areas regardless of task significance. The Adaptive Boosting meta-algorithm, which won the 2003 Gödel Prize, analytically assigns higher weights to misclassified cases to reclassify. However, it relies on weaker base classifiers that are iteratively strengthened, limiting improvements from base classifiers. Combining visual and computational approaches enables selecting stronger base classifiers before boosting. This paper proposes moving boosting methodology from focusing on only misclassified cases to all cases in the class overlap areas using Computational and Interactive Visual Learning (CIVL) with a Human-in-the-Loop. It builds classifiers in lossless visualizations integrating human domain expertise and visual insights. A Divide and Classify process splits cases to simple and complex, classifying these individually through computational analysis and data visualization with lossless visualization spaces of Parallel Coordinates or other General Line Coordinates. After finding pure and overlap class areas simple cases in pure areas are classified, generating interpretable sub-models like decision rules in Propositional and First-order Logics. Only multidimensional cases in the overlap areas are losslessly visualized simplifying end-user cognitive tasks to identify difficult case patterns, including engineering features to form new classifiable patterns. Demonstration shows a perfectly accurate and losslessly interpretable model of the Iris dataset, and simulated data shows generalized benefits to accuracy and interpretability of models, increasing end-user confidence in discovered models.

人机协作可解释性视觉分析分类模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。