arXiv:2511.20586cs.AIcs.LG2025-11

用主观逻辑构建神经网络信任传播框架,提升模型在对抗环境下的可靠性评估

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

  • 通过信任节点与信任函数并行计算输入、参数和激活的信任度
  • 在真实和对抗数据集上实现可解释的对称信任值,识别出高置信但不可靠的预测
  • 适合关注AI可信性、安全性与鲁棒性的研究人员和工程应用

可信性已成为人工智能系统在安全关键应用中部署的关键要求。传统评价指标如准确率和精确率无法有效捕捉模型预测中的不确定性或可靠性,尤其在对抗或退化条件下。本文提出并行信任评估系统(PaTAS),一种基于主观逻辑(Subjective Logic, SL)建模与传播神经网络信任的框架。PaTAS通过信任节点和信任函数与标准神经计算并行运行,实现输入、参数和激活信任在全网络中的传播。该框架定义了参数信任更新机制以在训练中优化参数可靠性,并提出推理路径信任评估(IPTA)方法,在推理阶段生成实例级信任值。在真实世界与对抗数据集上的实验表明,PaTAS生成可解释、对称且收敛的信任估计,补充了准确率指标,揭示了中毒、偏斜或不确定数据场景下的可靠性差距。结果证明,PaTAS能有效区分良性与对抗输入,并识别模型置信度与实际可靠性相悖的情况。通过在神经架构中实现透明且可量化的信任推理,PaTAS为整个AI生命周期的模型可靠性评估提供了基础。

原文摘要 · Abstract (English)

Trustworthiness has become a key requirement for the deployment of artificial intelligence systems in safety-critical applications. Conventional evaluation metrics, such as accuracy and precision, fail to appropriately capture uncertainty or the reliability of model predictions, particularly under adversarial or degraded conditions. This paper introduces the Parallel Trust Assessment System (PaTAS), a framework for modeling and propagating trust in neural networks using Subjective Logic (SL). PaTAS operates in parallel with standard neural computation through Trust Nodes and Trust Functions that propagate input, parameter, and activation trust across the network. The framework defines a Parameter Trust Update mechanism to refine parameter reliability during training and an Inference-Path Trust Assessment (IPTA) method to compute instance-specific trust at inference. Experiments on real-world and adversarial datasets demonstrate that PaTAS produces interpretable, symmetric, and convergent trust estimates that complement accuracy and expose reliability gaps in poisoned, biased, or uncertain data scenarios. The results show that PaTAS effectively distinguishes between benign and adversarial inputs and identifies cases where model confidence diverges from actual reliability. By enabling transparent and quantifiable trust reasoning within neural architectures, PaTAS provides a foundation for evaluating model reliability across the AI lifecycle.

信任评估主观逻辑神经网络鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。