用拓扑距离追踪日志对大模型输出的影响,提升安全决策可信度。
Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis

- 基于拓扑学设计段级归因距离,量化日志对模型输出几何结构的影响。
- 在真实攻防场景中,能自适应定位最关键影响日志,准确率显著提升。
- 适合需要可解释性与证据溯源的网络安全与智能体系统应用。
大语言模型(LLMs)正被广泛应用于网络安全领域,辅助分析师快速应对新兴威胁。然而,在安全操作中使用LLM的关键挑战在于输出可信度。随着智能体(Agentic AI)融入运营系统,必须具备可靠的证据归因与溯源能力,以追踪生成结果的来源。当自主代理做出决策时,能否回溯决策链至关重要,否则无法确定是哪一段数据导致了模型输出。现有方法难以区分复杂且高度相似的证据源,如网络攻击日志,暴露出核心缺陷:无法有效捕捉检索证据与生成响应之间的全局几何关系,影响证据验证可靠性。为此,本文提出拓扑归因距离(Topological Attribution Distance, TAD),受拓扑学启发,用于刻画输出及其与检索日志间几何形状的全局变化。若某条日志的嵌入向量在嵌入空间中显著改变模型输出的几何结构,则表明该日志是生成结果的关键来源。TAD通过段级消融归因方法,分析真实攻击事件中的日志,实现对关键日志的自适应定位。该方法为每一步隐藏状态提供可解释的追溯路径,使模型生成过程在几何层面更具可解释性,并支持网络安全和智能体工作流中的证据验证。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decision-making against emerging threats. However, there is a main criteria that must be met when using LLMs in cybersecurity, that is, trust in the generated outputs. As Agentic AI is integrated into operational systems, a robust evidence attribution and provenance tracking technique is essential to trace the origins of model generations. When autonomous agents make a decision (right or wrong), the ability to trace back through the decision chain is critical, as without it, teams cannot identify which segment of the data caused the model generation. Existing methods often struggle to distinguish among complex and highly similar evidence sources, such as cyber incident logs. This reveals a key gap: current approaches do not adequately capture the holistic geometric relationship between the retrieved evidence and the generated response for reliable evidence verification. To bridge this gap, we propose Topological Attribution Distance (TAD), inspired by Topology, to characterize and capture the global geometric shape of an output and its changes against its retrieved logs. In other words, if the embeddings of a specific source log drastically changes the geometry of the model's response in the embedding space, this suggests that such log is a critical source for the model's generated response. Therefore, TAD is powered by segment-level ablation attribution to investigate incident logs of an actual cyberattack. We demonstrate how TAD finds the most attributed logs on LLM outputs in an adaptive manner. This can provide an explainable and trustworthy tracing based on each LLM's hidden state to understand how geometrically different retrieved logs influence the model generation, and provide evidence verification in cybersecurity and Agentic-AI workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。