解释可能永远无法完成,因此信任反而是必然的。
Why Trust in AI May Be Inevitable
- 将解释视为在知识网络中寻找连接路径的搜索过程。
- 即使理性诚实且知识重叠,解释仍可能因时间不足而失败。
- 适合关注AI可信性与人类信任机制的研究者阅读。
在人机交互中,解释通常被视为建立对AI系统信任的必要条件。但我们认为,信任反而可能是前提,因为解释有时根本不可能实现。基于将解释形式化为在知识网络中搜索连接路径的过程,我们发现:即使交互双方理性、诚实、沟通无误且知识重叠,解释仍可能失败。这是因为成功解释不仅需要共享知识的存在,还需在有限时间内找到连接路径;当时间耗尽时,理性行为是停止尝试,而非继续搜索。这一发现对人机交互具有重要意义:随着大型语言模型越来越擅长生成表面合理但虚假的解释,人类可能倾向于默认信任,而非追求真实解释。这带来了错误信任和知识整合不充分的风险。
原文摘要 · Abstract (English)
In human-AI interactions, explanation is widely seen as necessary for enabling trust in AI systems. We argue that trust, however, may be a pre-requisite because explanation is sometimes impossible. We derive this result from a formalization of explanation as a search process through knowledge networks, where explainers must find paths between shared concepts and the concept to be explained, within finite time. Our model reveals that explanation can fail even under theoretically ideal conditions - when actors are rational, honest, motivated, can communicate perfectly, and possess overlapping knowledge. This is because successful explanation requires not just the existence of shared knowledge but also finding the connection path within time constraints, and it can therefore be rational to cease attempts at explanation before the shared knowledge is discovered. This result has important implications for human-AI interaction: as AI systems, particularly Large Language Models, become more sophisticated and able to generate superficially compelling but spurious explanations, humans may default to trust rather than demand genuine explanations. This creates risks of both misplaced trust and imperfect knowledge integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。