让AI像医生一样有依据地诊断,还能验证其推理是否靠谱。
Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
- 用强化学习训练AI主动寻找外部视觉证据支持判断
- 比非交互式基线模型准确率提升,Brier得分降低18%
- 通过遮蔽证据测试,证明其推理依赖真实证据
高风险领域如医疗中的AI解释常缺乏可验证性,影响可信度。为此,我们提出一种交互式智能体,通过可审计的动作序列生成解释。该智能体学习策略性地获取外部视觉证据以支持诊断推理,采用强化学习优化,使模型兼具高效性与泛化能力。实验表明,这种基于动作的推理过程显著提升校准准确率,相比非交互基线,Brier得分降低18%。为验证解释的忠实性,引入因果干预方法:遮蔽智能体选择的视觉证据后,性能明显下降(ΔBrier=+0.029),证实证据对其决策至关重要。本工作为构建可验证、忠实推理的AI系统提供了实用框架。
原文摘要 · Abstract (English)
Explanations for AI models in high-stakes domains like medicine often lack verifiability, which can hinder trust. To address this, we propose an interactive agent that produces explanations through an auditable sequence of actions. The agent learns a policy to strategically seek external visual evidence to support its diagnostic reasoning. This policy is optimized using reinforcement learning, resulting in a model that is both efficient and generalizable. Our experiments show that this action-based reasoning process significantly improves calibrated accuracy, reducing the Brier score by 18\% compared to a non-interactive baseline. To validate the faithfulness of the agent's explanations, we introduce a causal intervention method. By masking the visual evidence the agent chooses to use, we observe a measurable degradation in its performance ($Δ$Brier=+0.029), confirming that the evidence is integral to its decision-making process. Our work provides a practical framework for building AI systems with verifiable and faithful reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。