用数学方法精确追踪神经网络决策背后的训练案例支持
From Neural Network Decisions to Training Cases: An Exact Account via Case-Based Decision Theory

- 通过最小二乘法构建顶层探测器,实现决策得分的案例分解
- 在多个数据集上恢复案例偏好结构,Top-30一致性达最优
- 无需重训模型或访问训练轨迹,适合高风险场景审计
神经网络在医疗诊断、信贷审批等高风险领域日益关键,其决策需提供案例级证据:哪些训练样本支持当前动作及其对应结果。本文基于案例决策理论(CBDT),证明固定神经表示下的最小二乘法(OLS)动作读出可实现精确的案例分解。每个动作得分是训练案例回报的加权和,权重由经验格拉姆几何决定。识别出满足CBDT相似性语义的充分条件;超出该条件时,权重应视为符号化的格拉姆几何影响。该分解能生成审计信号,追溯得分来源、衡量决策一致性并识别弱支持。在合成CBDT、PJM、Adult Income和Default Credit任务中,该方法成功恢复案例偏好结构,在对比归因基线中取得最高平均Top-30一致性,同时保持支持重构竞争力。审计仅需拟合一个OLS顶层探测器,无需重新训练表示或访问原始优化轨迹;探测器保真度通过得分重建效果衡量。
原文摘要 · Abstract (English)
Neural networks increasingly guide decisions in high-stakes domains such as medical diagnosis, credit approval, and energy bidding. Audit in these settings requires case-level evidence: which training cases support an action and what outcomes they carried. Case-based decision theory (CBDT) formalizes this reasoning by aggregating outcome support from remembered cases. We show that an OLS action readout fitted on a fixed neural representation admits an exact case-based decomposition. Each action score is a weighted sum of training-case returns, with coefficients determined by empirical Gram geometry. We identify a sufficient regime for CBDT similarity semantics; outside it, the coefficients should generally be treated as signed Gram-geometric influence. The decomposition yields audit signals that trace scores to training cases, measure action coherence, and identify weak support. Across synthetic CBDT, PJM, Adult Income, and Default Credit tasks, the method recovers case-level preference structure and achieves the highest mean Top-30 consistency among compared attribution baselines, while remaining competitive on support reconstruction. The audit requires only fitting an OLS top-layer probe, without retraining the representation or accessing the original optimization trajectory; probe fidelity is measured by score reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。