arXiv:2512.00015cs.HCcs.AI2025-12中稿 · The World Conferen…

用概念模型提升人机协作可解释性,但效果受双方认知差异制约。

The Impact of Concept Explanations and Interventions on Human-Machine Collaboration

  • 通过概念瓶颈模型显式建模人类定义的概念,增强决策可解释性。
  • 人机协同中模型干预虽提升可解释性,但未显著提高任务准确率。
  • 需多次交互才能理解模型逻辑,认知错位会削弱协作效果。

深度神经网络因决策过程不透明常被视为黑箱。为降低这种不透明性,研究引入概念模型(CMs),如概念瓶颈模型(CBMs),通过预测人类定义的概念作为中间步骤来提升可解释性。在人机协作场景中,更高的可解释性有助于人类理解并信任模型。早期研究表明,当模型对概念的错误预测被真实值替代(即概念干预)时,任务准确率会提升。然而,该结果未经过人类评估,因此其在人机协同中的适用性仍未知。本文首次在真实人机协作设置下开展实验,使用CBMs进行人类研究。结果表明,与标准DNN相比,CBMs显著提升了可解释性,促进了人机对齐;但这种对齐并未带来任务准确率的显著提升。理解模型决策需要多次交互,而模型与人类决策过程的错位可能损害可解释性及模型有效性。

原文摘要 · Abstract (English)

Deep Neural Networks (DNNs) are often considered black boxes due to their opaque decision-making processes. To reduce their opacity Concept Models (CMs), such as Concept Bottleneck Models (CBMs), were introduced to predict human-defined concepts as an intermediate step before predicting task labels. This enhances the interpretability of DNNs. In a human-machine setting greater interpretability enables humans to improve their understanding and build trust in a DNN. In the introduction of CBMs, the models demonstrated increased task accuracy as incorrect concept predictions were replaced with their ground truth values, known as intervening on the concept predictions. In a collaborative setting, if the model task accuracy improves from interventions, trust in a model and the human-machine task accuracy may increase. However, the result showing an increase in model task accuracy was produced without human evaluation and thus it remains unknown if the findings can be applied in a collaborative setting. In this paper, we ran the first human studies using CBMs to evaluate their human interaction in collaborative task settings. Our findings show that CBMs improve interpretability compared to standard DNNs, leading to increased human-machine alignment. However, this increased alignment did not translate to a significant increase in task accuracy. Understanding the model's decision-making process required multiple interactions, and misalignment between the model's and human decision-making processes could undermine interpretability and model effectiveness.

人机协作可解释性概念模型认知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。