arXiv:2606.06081cs.AIcs.HC2026-06

首个衡量人类对集合型AI建议合理依赖的框架。

A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice

论文配图:A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice
图 1 · 摘自论文原文
  • 提出分类与回归任务中评估集合型AI建议的新指标。
  • 区分依赖行为与依赖质量,揭示人类决策中的细微差异。
  • 适合研究人机协作、不确定性表达的学者参考。

在人机协同中,对AI建议的合理依赖已成为核心研究课题。现有框架仅关注点预测型建议,但集合型建议(如离散集合或连续区间)正被广泛用于表达不确定性并提升人类决策。本文首次构建了在序列判断-建议范式下,针对分类与回归任务的集合型AI建议合理依赖的正式评估框架。对于分类任务,引入评价集合型建议的维度,并定义‘正确依赖AI率’与‘正确依赖自我率’,共同刻画合理依赖。对于回归任务,提出‘AI依赖量’与‘AI依赖质量’,分别衡量决策者是否使用了AI建议,以及依赖是否使其更接近真实值。通过该框架的应用,我们展示了这些指标如何捕捉现有度量所忽视的人机协作关键细节。

原文摘要 · Abstract (English)

Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusively on point predictions as AI advice. However, set-valued AI advice (e.g., discrete sets or continuous intervals) is increasingly being used to communicate uncertainty and improve human decision making. In this paper, we develop the first formal framework for measuring appropriate reliance on set-valued AI advice within the sequential judge-advisor paradigm, spanning both classification and regression tasks. For classification, we first introduce the dimensions that are necessary for evaluating set-valued AI advice. We then define two metrics: correct reliance rate on AI and correct reliance rate on self, which jointly characterize appropriate reliance in this setting. For regression, we introduce quantity of AI reliance and quality of AI reliance, which respectively measure whether a decision maker utilized the AI advice and whether their reliance helped them get closer to the ground truth relative to their initial estimate. Through the application of our framework, we demonstrate how these metrics capture important nuances in human-AI collaboration that existing measures overlook.

人机协同AI建议不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。