arXiv:2605.12646cs.LGcs.AI2026-05被引 1

AI与人类信心对齐能显著降低决策学习难度。

Learning to Decide with AI Assistance under Human-Alignment

  • 将人机决策问题转化为上下文在线学习模型,分析信心对齐的影响。
  • 理论证明信心对齐可将学习误差降至O(√(T log T)),优于无对齐情况。
  • 实验验证理论在真实人类行为中依然稳健,即使对齐不完美也有效。

当人工智能在高风险领域辅助决策时,其预测置信度的传达至关重要。然而实证表明,决策者难以仅凭该置信度判断是否信任预测。近期研究发现,人工智能与人类信心的对齐程度与辅助决策效用正相关。本文聚焦二元预测与二元决策场景,证明该问题等价于具有完整反馈的两臂在线上下文学习问题,并建立期望后悔下界Ω(√(|H|·|B|·T)),其中H、B分别为人类与AI的置信值集合。当两者完全对齐时,学习者可实现期望后悔O(√(|H|·T log T));在√|H|=O(log T)且B可数条件下,利用广义Dvoretzky-Kiefer-Wolfowitz不等式进一步将后悔降至O(√(T log T))。实验基于两项真实人类受试研究数据,验证了理论结果在信心不对齐情况下的鲁棒性。

原文摘要 · Abstract (English)

It is widely agreed that when AI models assist decision-makers in high-stakes domains by predicting an outcome of interest, they should communicate the confidence of their predictions. However, empirical evidence suggests that decision-makers often struggle to determine when to trust a prediction based solely on this communicated confidence. In this context, recent theoretical and empirical work suggests a positive correlation between the utility of AI-assisted decision-making and the degree of alignment between the AI confidence and the decision-makers' confidence in their own predictions. Crucially, these findings do not yet elucidate the extent to which this alignment influences the complexity of learning to make optimal decisions through repeated interactions. In this paper, we address this question in the canonical case of binary predictions and binary decisions. We first show that this problem is equivalent to a two-armed online contextual learning problem with full feedback, and establish a lower bound of $Ω(\sqrt{|H| \cdot |B| \cdot T} )$ on the expected regret any learner can attain, where $H$ and $B$ denote the sets of human and AI confidence values. We then demonstrate that, under perfect alignment between AI and human confidence, a learner can attain an expected regret of $O(\sqrt{|H| \cdot T\log T})$ and, when $\sqrt{|H|} = O(\log T)$ and $B$ is countable, a non-trivial generalization of the Dvoretzky-Kiefer-Wolfowitz inequality improves the regret bound to $O(\sqrt{T\log T})$. Taken together, these results reveal that alignment can reduce the complexity of learning to make decisions with AI assistance. Experiments on real data from two different human-subject studies where participants solve simple decision-making tasks assisted by AI models show that our theoretical results are robust to violations of perfect alignment.

人机协同决策学习置信度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。