为高风险场景下的生成式AI建立可信依赖的评估框架。
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows
- 提出三个必要条件:认知谦逊、可访问性与反认知不公。
- 案例表明现有指标无法捕捉系统性认知危害。
- 适合关注AI伦理与人机协作的研究者与设计者。
生成式AI正被广泛应用于法律推理、医疗决策和招聘等高风险专业场景,其输出影响用户信念、推理过程与认知判断。当前主流框架关注准确性、公平性、可解释性或用户信任,但未明确界定‘合理依赖’这一核心评价目标——即用户在何种条件下可正当将AI输出作为自身推理依据。本文基于哲学中的信任理论,提出‘认知可信性’概念,构建一个包含三项共同必要且不可替代条件的规范框架:第一,认知谦逊要求系统清晰表达自身能力边界;第二,认知可及性要求用户能情境化地审查、质疑和反驳输出;第三,抵抗认知不公要求系统承认用户为合法认知主体,避免边缘化其知识与经验。通过真实案例分析,我们揭示了上述三类失效会引发标准指标无法捕捉的严重后果。最后,本文提出以‘认知合理依赖’为核心的设计与评估路径。
原文摘要 · Abstract (English)
Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。