arXiv:2604.03237cs.HCcs.AI2026-04

不同任务需匹配不同AI辅助方式,才能实现人机协作的精准依赖。

Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks

  • 针对视觉与语言任务设计差异化的AI支持策略
  • 解释类支持在逻辑推理中提升准确率,但对视觉任务无效
  • 预测概率最利于用户识别错误并纠正

随着AI系统越来越多地支持人类决策,关键挑战在于如何提供信息,使人们能判断何时应信赖AI预测、何时应质疑或否决。通过三项控制性实验,分别考察了在RAVEN矩阵(抽象视觉推理)和LSAT问题(演绎逻辑推理)中的表现。多阶段揭示研究显示,AI预测与解释对客观准确率和主观信心的影响不同。在视觉推理中,大模型(LLM)解释未能提升准确率,仅提供预测结果已足够,且预测概率在描述准确性和错误恢复上表现最佳;基于概率的自适应自动化策略提供了更优性能基准。而在语言逻辑推理中,LLM解释显著优于专家撰写解释与概率支持,带来最高准确率与错误恢复能力。结果表明,不存在通用有效的支持策略。人机界面应根据任务类型与用户可获取证据,设计支持校准依赖与高效纠错的辅助机制。

原文摘要 · Abstract (English)

As AI systems increasingly support human decision making, a central challenge is determining what information helps people recognize when to rely on AI predictions and when to question or override them. Across three controlled human-subject studies spanning abstract visual reasoning with RAVEN matrices and deductive logical reasoning with LSAT problems, we examine how different forms of AI support affect human--AI team performance. A multi-stage reveal study shows that AI predictions and explanations can affect objective accuracy and subjective confidence differently. In visual reasoning, LLM explanations do not improve accuracy beyond the predicted answer alone, and no additional support format significantly outperforms prediction-only support; predicted probabilities show the highest descriptive accuracy and error recovery, while a derived selective-automation policy provides a higher-performing reference benchmark. In language-based logical reasoning, by contrast, LLM explanations yield the highest accuracy and error recovery, outperforming expert-written explanations and probability-based support. These results show that no single support strategy is universally effective. Human--AI interfaces should instead be designed to support calibrated reliance and effective error recovery by matching the form of assistance to the task and the evidence available to users.

人机协作可信AI决策支持提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。