arXiv:2608.14352cs.SEcs.LG2026-08中稿 · publication at ACM…

用自动化方法从大模型智能体行为中提炼可解释的策略模型

ATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata Learning

  • 通过轨迹抽象与自动机学习,将复杂行为转化为有限状态模型
  • 在渗透测试场景中识别出12台脆弱机器的高阶攻击策略
  • 适合关注AI可解释性、安全审计与模型压缩的研究者

基于大语言模型(LLM)的智能体在软件测试和网络安全评估等复杂任务中日益普及。然而其行为难以理解与分析。现有评估主要关注任务成功率和执行轨迹,难以揭示智能体的实际策略。本文提出ATLAS(Automata Learning for Agent Trajectory Analysis and Strategy Discovery),一种从智能体轨迹中恢复可解释行为模型的方法。ATLAS结合轨迹抽象与自动机学习,推断出捕捉智能体-环境交互策略的有限状态模型。这些模型提供人类可读的洞察,支持对重复行为、决策点、成功路径及失败循环的自动化分析。作为概念验证,我们将ATLAS应用于基于LLM的渗透测试智能体生成的轨迹,结果揭示了从原始轨迹中难以察觉的高阶攻击策略。我们进一步展示了如何利用学习到的行为模型支持可解释性、模型引导探索、审计与系统分析,并实现从前沿大模型到小型语言模型的知识迁移。在包含12个脆弱机器的渗透测试案例中,模型变换生成了简洁的行为解释。ATLAS为模型驱动工程开辟新路径:将智能体轨迹转化为显式行为模型,实现对原本不透明的AI智能体的系统性理解与分析。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based agents are increasingly used for complex tasks such as software testing and cybersecurity assessment. While these agents demonstrate impressive capabilities, their behavior is difficult to understand, explain, and analyze. Existing evaluations focus mainly on task success and execution traces, offering limited insight into the strategies employed by the agent. We present ATLAS (Automata Learning for Agent Trajectory Analysis and Strategy Discovery), an approach for recovering interpretable behavioral models from agent trajectories. ATLAS combines trace abstraction with automata learning to infer finite-state models that capture observed agent-environment interaction strategies. These models provide human-interpretable insights and support automated analyses of recurring behaviors, decision points, successful task-completion paths, and failure loops. As a proof of concept, we apply ATLAS to trajectories generated by an LLM-based penetration-testing agent. The resulting models expose high-level behavioral strategies for exploiting vulnerable machines that are difficult to identify from raw execution traces alone. We discuss how learned behavioral models can support explainability, model-guided exploration, auditing, and analysis of agentic systems. We further demonstrate symbolic model-based knowledge transfer from powerful frontier models to compact language models. In addition, we show how model transformations can derive concise explanations of agent behavior in a penetration-testing case study comprising 12 vulnerable machines. ATLAS highlights a new opportunity for model-driven engineering: transforming agent trajectories into explicit behavioral models that enable systematic understanding and analysis of otherwise opaque AI agents.

智能体行为可解释性自动机学习渗透测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。