arXiv:2605.20439cs.LGcs.HC2026-05中稿 · Thirty-Fourth Euro…

对话式可解释AI能否提升用户表现?实验发现它能但效果不显著。

Can Conversational XAI Improve User Performance? An Experimental Study

  • 设计实验对比对话与问答式解释,评估用户预测性能提升。
  • 42名参与者均超越模型表现,但两种解释方式无明显差异。
  • 适合关注人机协作与解释交互设计的研究者参考。

可解释人工智能(XAI)旨在揭示预测模型的内在机制并提升用户表现,但实际效果常不达预期。对话式XAI助手有望克服这些局限,但其对客观性能指标的影响仍缺乏实证支持。本文提出一种实验设计,通过预测准确率、模型理解力和错误识别能力评估解释辅助的效果。采用可解释性内置的预测模型,构建用户可通过识别并纠正系统性错误超越模型的情境。比较对话式辅助与问答式辅助在支持用户理解模型解释方面的表现。初步测试结果(N=42)显示,两组参与者均显著优于模型,但两种辅助方式在性能上无显著差异,且整体参与度较低。这些发现为后续全规模研究提供了优化方向,包括增强用户参与度的干预措施及对性能提升机制的深入探究。

原文摘要 · Abstract (English)

Explainable AI (XAI) techniques aim to provide insights into predictive models and enhance user performance, yet they often fall short of these expectations. Conversational XAI assistants promise to overcome such limitations, but empirical evidence on their impact on objective performance measures remains limited. We propose an experimental design for evaluating explanation assistance through prediction accuracy, model understanding, and error identification. Using an explainable-by-design prediction model, we create conditions where users can outperform the model by identifying and compensating for systematic errors. We compare conversational assistance against Q&A-based assistance to assess which better supports users in working with model explanations. Preliminary results from testing our experimental design show that participants (N=42) in both treatments significantly outperformed the model but reveal no performance differences between assistance types and modest engagement overall. These findings inform refinements for our planned full study, including enhanced engagement interventions and investigation of the mechanisms driving improved predictions.

可解释AI人机交互实验研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。